Related Experiment Video
Updated: Sep 13, 2025

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
Published on: October 27, 2023
Natural language processing improves reliable identification of COVID-19 compared to diagnostic codes alone
Nathaniel Hendrix1, Rishi V Parikh2, Madeline Taskier1
1Center for Professionalism and Value in Health Care, American Board of Family Medicine, Washington, DC 20036, United States.
Diagnostic codes for COVID-19 have limited accuracy, especially in older adults and Black patients. Natural language processing (NLP) shows promise but requires frequent updates for reliable patient cohort identification.
Area of Science:
- Health Informatics
- Medical Record Analysis
- Epidemiological Research
Background:
- Observational studies on COVID-19 frequently utilize diagnostic codes for patient classification.
- The accuracy of these codes and potential for differential misclassification across diverse patient demographics remain largely unexamined.
- Understanding these limitations is crucial for reliable real-world evidence generation.
Purpose of the Study:
- To evaluate the accuracy of diagnostic codes for identifying COVID-19 patients.
- To investigate age, race, and ethnicity as predictors of differential misclassification.
- To compare the performance of diagnostic codes against natural language processing (NLP) classifiers.
Main Methods:
- Proof-of-concept study using two primary care cohorts from the American Family Cohort.
- Comparison of ICD-10 diagnostic codes against NLP classifiers trained on physician-assessed clinical notes.
- Three NLP models (tree-based, recurrent neural network, transformer-based) were trained and tested.
Main Results:
- Only 63% of likely COVID-19 patients had a documented ICD-10 code.
- Code sensitivity was lower in older patients (60.6% for 75+) and Black patients (58.5%).
- A tree-based NLP classifier achieved an AUC of 0.92 but showed decreased accuracy in older patients and required frequent retraining.
Conclusions:
- Diagnostic codes exhibit significant limitations in accurately identifying COVID-19 patients, with notable disparities across age, race, and ethnicity.
- NLP holds potential for improving cohort identification but necessitates continuous retraining to maintain performance.
- Future research should focus on robust methods for accurate patient classification in observational studies.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Single Nucleotide Polymorphisms-SNPs
Methods of Classification and Identification
Improving Translational Accuracy
MALDI-TOF Mass Spectrometry
Matrix-assisted laser desorption ionization (MALDI) is a commonly...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...

