Related Experiment Video
Updated: Aug 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
MLM-based typographical error correction of unstructured medical texts for named entity recognition.
Eun Byul Lee1, Go Eun Heo2, Chang Min Choi3
1Department of Digital Analytics, Yonsei University, 50 Yonsei-ro Seodaemun-gu, 03722, Seoul, Republic of Korea.
This study introduces a novel typo correction model for medical texts, significantly improving information extraction accuracy. The model enhances natural language processing tasks by addressing errors in electronic health records.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Computational Linguistics
Background:
- Electronic Health Records (EHRs) contain valuable research data but are hindered by frequent typographical errors.
- Extracting structured information from unstructured medical text is challenging due to data quality issues.
- Existing methods for correcting typos in medical records are limited, necessitating improved approaches.
Purpose of the Study:
- To develop and evaluate a novel typo correction model for unstructured medical texts.
- To enhance the accuracy of information extraction from surgical pathology records and other medical data.
- To overcome the limitations of current methods in handling real-world medical data errors.
Main Methods:
- A context-aware typo correction model based on the Masked Language Model (MLM) was developed.
- A word dictionary derived from PubMed abstracts was integrated into the model.
- Fine-tuning of a pre-trained BERT model followed by deep learning-based Named Entity Recognition (NER) was performed on corrected data.
Main Results:
- The proposed typo correction model demonstrated superior performance compared to the existing SymSpell model.
- The model achieved an approximate 5% and 9% F1-score improvement on the NCBI-disease and surgical pathology datasets, respectively.
- Named Entity Recognition (NER) performance increased by 2% on the NCBI-disease dataset and showed a significant 25% improvement on surgical pathology records post-correction.
Conclusions:
- Typographical errors in unstructured medical text significantly degrade natural language processing task performance.
- The proposed context-aware typo correction model offers a robust solution for real-world medical data analysis.
- This methodology effectively addresses data quality challenges, improving the reliability of information extraction from medical records.
Related Concept Videos
Types of Errors: Detection and Minimization
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Mismatch Repair
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Improving Translational Accuracy
Genetic Lingo
Detection of Gross Error: The Q Test

