A machine learning-based phenotype for long COVID in children: an EHR-based study from the RECOVER program

Vitaly Lorman1, Hanieh Razzaghi1, Xing Song2

  • 1Applied Clinical Research Center, Children's Hospital of Philadelphia, Philadelphia, PA, United States.

Insights

A machine learning algorithm was developed to identify pediatric Post-Acute Sequelae of SARS CoV-2 (PASC) in electronic health records. This tool aids in classifying PASC cases, distinguishing between MIS-C and non-MIS-C variants for research and clinical trial recruitment.

Area of Science:

  • Pediatric Health Informatics
  • Machine Learning in Clinical Research
  • Epidemiology of Post-Acute Sequelae of SARS CoV-2 (PASC)

Background:

  • Clinical understanding and definitions of pediatric Post-Acute Sequelae of SARS CoV-2 (PASC) are evolving.
  • Reliable identification of PASC patients within health systems data is crucial for research and clinical care.
  • Distinguishing PASC from Multisystem Inflammatory Syndrome in Children (MIS-C) is important for accurate classification.

Approach:

  • Developed and validated a machine learning algorithm using the PEDSnet Electronic Health Record (EHR) network.
  • Selected patient features from conditions, procedures, diagnostics, and medications using a tree-based scan statistic.
  • Employed an XGBoost model with hyperparameter tuning via cross-validated grid search and evaluated using 5-fold cross-validation.
  • Utilized Shapley Additive exPlanations (SHAP) for model prediction and feature importance analysis.

Key Points:

  • The algorithm effectively classifies pediatric patients with PASC from electronic health records.
  • Feature importance analysis using SHAP values provides insights into characteristics associated with PASC.
  • The model demonstrated robust performance through rigorous validation methods.

Conclusions:

  • The developed machine learning model serves as a valuable tool for identifying and characterizing PASC in pediatric populations.
  • The model's flexibility allows for precise patient identification for studies or broad screening for clinical trials.
  • This approach enhances the utility of health systems data for PASC research, especially where diagnostic codes are unreliable.
Abstract