Related Experiment Videos
Detecting Misspelled Drug Names Using Transformer-Based Language Models: Model Development and External Validation
Jiayu Lu1, Kevin W McConeghy2,3, Andrew R Zullo2,3,4
1Department of Medicine/Section of Preventive Medicine and Epidemiology, Boston University Chobanian & Avedisian School of Medicine, Boston, MA, United States.
Background:
Misspellings in medication names can compromise patient safety, reduce data utility, and impede large-scale data initiatives that integrate medication information from electronic health records (EHRs). Existing methods for detecting misspelled medical terms are mostly dictionary-based and can lead to high false-positive rates when correctly spelled but previously unseen (out-of-vocabulary) terms are encountered.
Objective:
We aimed to develop and validate domain-specific, transformer-based language models for detecting misspelled drug names, with an emphasis on performance for unseen terms.
Methods:
Using RxNorm as a standardized drug vocabulary, we created an RxNorm-augmented training corpus and developed two BERT (Bidirectional Encoder Representations from Transformers)-based models-BERTDrug and CharBERTDrug-for misspelling detection. Specifically, we randomly split 69,824 RxNorm drug names into training, development, and test sets (3:1:1) and generated k misspellings per name using text-perturbation techniques (k optimized for training; fixed at 1 for development and test sets). The models were fine-tuned on the training and development sets and evaluated using the RxNorm test set and 3586 drug names from the Long-Term Care Data Cooperative (LTCDC) database (external validation). The RxNorm test set and out-of-vocabulary LTCDC dataset (1922 terms), neither overlapping with the RxNorm training data, were used to evaluate performance on unseen terms. SpellChecker served as a dictionary-based baseline, while fastTextML and BioWordVecML, which used different subword embeddings as inputs for machine learning, served as additional baselines. Additionally, we compared model performance with GPT-4o, a generative large language model (LLM), using 2200 randomly sampled test terms.
Results:
On the RxNorm test set, BERTDrug and CharBERTDrug outperformed the baseline models across most metrics. BERTDrug achieved the best overall performance (F1-score=0.859; area under the receiver operating characteristic curve [ROC-AUC]=0.947), followed by CharBERTDrug (F1-score=0.833; ROC-AUC=0.906). Both models also outperformed the baseline models on the out-of-vocabulary LTCDC dataset across most metrics, with CharBERTDrug performing best (F1-score=0.696; ROC-AUC=0.788), followed by BERTDrug (F1-score=0.669; ROC-AUC=0.786). In the secondary analysis, both models exceeded GPT-4o on most metrics (except Recall) for 2000 RxNorm terms. BERTDrug performed best (ROC-AUC=0.951; F1-score=0.855), followed by CharBERTDrug (ROC-AUC=0.911; F1-score=0.831) and GPT-4o (ROC-AUC=0.856; F1-score=0.721). In contrast, among 200 randomly selected LTCDC terms, GPT-4o performed best on most metrics except precision and specificity.
Conclusions:
Domain-specific language models improved detection of misspellings in out-of-vocabulary drug names and outperformed baseline models in both internal and external evaluations. The comparison with a generative LLM suggests that domain shift may substantially reduce the advantages conferred by domain-specific training. With further fine-tuning on diverse data that capture the terminology, formatting conventions, and spelling patterns encountered across real-world clinical settings, these models could be adapted for use in other clinical databases and EHR systems to improve medication data quality for research and to support future safety-focused applications.