Related Experiment Video
Updated: Jun 24, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
EHR-based Case Identification of Pediatric Long COVID: A Report from the RECOVER EHR Cohort
Morgan Botdorf1, Kimberley Dickinson1, Vitaly Lorman1
1Applied Clinical Research Center, Children's Hospital of Philadelphia, Philadelphia, PA.
Insights
A new algorithm for identifying pediatric Long COVID showed moderate accuracy when compared to manual chart reviews. Adjusting the algorithm to account for pre-existing conditions improved its performance in identifying Long COVID cases in children.
Area of Science:
- Pediatric Health
- Infectious Diseases
- Clinical Informatics
Background:
- Long COVID presents persistent symptoms post-COVID-19 infection in children, but lacks a unified clinical definition.
- Understanding and identifying Long COVID in pediatric populations is crucial for research and clinical management.
- Existing diagnostic methods for Long COVID in children are limited, necessitating the development of reliable identification tools.
Purpose of the Study:
- To evaluate the performance of an empirically derived Long COVID case identification algorithm (computable phenotype) against manual chart review in a pediatric cohort.
- To assess the accuracy and agreement between an electronic health record (EHR)-based algorithm and clinician chart review for identifying pediatric Long COVID.
- To identify reasons for discrepancies between algorithmic and manual case identification and explore methods to improve algorithm performance.
Main Methods:
- An algorithm using diagnostic codes associated with Long COVID was applied to a large EHR database of pediatric patients with SARS-CoV-2 infection.
- A subset of patients (n=651) identified by the algorithm were compared against manual chart review to assess overlap and discordance.
- Reasons for disagreements were analyzed, and the algorithm's performance was re-evaluated after incorporating consideration of prior medical conditions.
Main Results:
- The algorithm demonstrated moderate overlap with manual chart review for Long COVID identification (accuracy=0.62, PPV=0.49, NPV=0.75).
- Discrepancies were largely due to clinicians attributing Long COVID-like symptoms to pre-existing conditions.
- Algorithm performance improved significantly when prior medical conditions were factored into the analysis (accuracy=0.71, PPV=0.65, NPV=0.74).
Conclusions:
- The moderate agreement between the algorithm and chart review highlights challenges stemming from the lack of a standardized Long COVID definition in children.
- The study underscores the importance of considering pre-existing conditions when developing and validating Long COVID classification algorithms.
- Careful consideration of the strengths and limitations of both algorithmic and manual methods is essential for accurate Long COVID case identification in pediatric research.
Objective:
Long COVID, marked by persistent, recurring, or new symptoms post-COVID-19 infection, impacts children's well-being yet lacks a unified clinical definition. This study evaluates the performance of an empirically derived Long COVID case identification algorithm, or computable phenotype, with manual chart review in a pediatric sample. This approach aims to facilitate large-scale research efforts to understand this condition better.
Methods:
The algorithm, composed of diagnostic codes empirically associated with Long COVID, was applied to a cohort of pediatric patients with SARS-CoV-2 infection in the RECOVER PCORnet EHR database. The algorithm classified 31,781 patients with conclusive, probable, or possible Long COVID and 307,686 patients without evidence of Long COVID. A chart review was performed on a subset of patients (n=651) to determine the overlap between the two methods. Instances of discordance were reviewed to understand the reasons for differences.
Results:
The sample comprised 651 pediatric patients (339 females, M = 10.10 years) across 16 hospital systems. Results showed moderate overlap between phenotype and chart review Long COVID identification (accuracy = 0.62, PPV = 0.49, NPV = 0.75); however, there were also numerous cases of disagreement. No notable differences were found when the analyses were stratified by age at infection or era of infection. Further examination of the discordant cases revealed that the most common cause of disagreement was the clinician reviewers' tendency to attribute Long COVID-like symptoms to prior medical conditions. The performance of the phenotype improved when prior medical conditions were considered (accuracy = 0.71, PPV = 0.65, NPV = 0.74).
Conclusions:
Although there was moderate overlap between the two methods, the discrepancies between the two sources are likely attributed to the lack of consensus on a Long COVID clinical definition. It is essential to consider the strengths and limitations of each method when developing Long COVID classification algorithms.
More Related Videos
08:13Development and Implementation of a Multi-Disciplinary Technology Enhanced Care Pathway for Youth and Adults with Concussion
Published on: January 20, 2019
10:02Event Related Potentials ERPs and other EEG Based Methods for Extracting Biomarkers of Brain Dysfunction: Examples from Pediatric Attention Deficit/Hyperactivity Disorder ADHD
Published on: March 12, 2020
Related Concept Videos
Methods of Documentation VII: EMR
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...