Related Experiment Video
Updated: Sep 22, 2025

10:26
Author Spotlight: A 3D Digital Model for the Diagnosis and Treatment of Pulmonary Nodules
Published on: May 19, 2023
2.1K
Identifying who has long COVID in the USA: a machine learning approach using N3C data
Emily R Pfaff1, Andrew T Girvin2, Tellen D Bennett3
1Department of Medicine, UNC Chapel Hill School of Medicine, Chapel Hill, NC, USA.
The Lancet. Digital Health
|May 19, 2022
Summary
Machine learning models accurately identify potential long COVID patients using electronic health records. This aids in clinical trial recruitment and patient care for the evolving condition.
Area of Science:
- Health Informatics
- Machine Learning in Medicine
- Public Health
Background:
- Long COVID, or post-acute sequelae of SARS-CoV-2 infection, presents heterogeneous symptoms challenging precise definition.
- Electronic health records (EHRs) are vital for understanding Long COVID, as part of the NIH's RECOVER Initiative.
- Accurate identification of Long COVID patients is crucial for research and clinical management.
Purpose of the Study:
- Develop and validate machine learning models to identify potential Long COVID cases from EHR data.
- Enable efficient patient identification for clinical trials and specialized care.
- Provide a scalable method for Long COVID case ascertainment.
Main Methods:
- Utilized the National COVID Cohort Collaborative (N3C) EHR repository for a large adult patient cohort (n=1,793,604).
- Trained XGBoost models using demographics, healthcare utilization, diagnoses, and medications from 97,995 COVID-19 patients and 597 Long COVID clinic patients.
- Validated models on data from a fourth site, assessing performance across different patient subgroups (all, hospitalized, non-hospitalized).
Main Results:
- Achieved high accuracy in identifying potential Long COVID patients: AUC of 0.92 (all), 0.90 (hospitalized), and 0.85 (non-hospitalized).
- Key predictive features included healthcare utilization rates, patient age, dyspnea, and other EHR-documented diagnoses and medications.
- Shapley values were used to determine feature importance for model interpretability.
Conclusions:
- The developed models serve as a valuable proxy for identifying patients needing Long COVID specialty care.
- Facilitates the urgent need to identify Long COVID patients for clinical trials.
- Models are adaptable and can be retrained with new data sources to improve accuracy and meet specific study requirements.

