Related Experiment Video
Updated: Jan 18, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
A natural language processing pipeline for identifying pediatric long COVID symptoms and functional impacts in
H Timothy Bunnell1, Cara Reedy1, Vitaly Lorman2
1Biomedical Research Informatics Center, Nemours Children's Health, Wilmington, DE 19803, United States.
Insights
A new natural language processing (NLP) pipeline accurately identifies Long COVID symptoms and functional impacts in children using electronic health records. This method enhances understanding of pediatric Long COVID, improving patient characterization.
Area of Science:
- Pediatric Health
- Medical Informatics
- Natural Language Processing
Background:
- Long COVID presents a significant challenge in pediatric populations, with symptoms and functional impacts often poorly characterized in electronic health records (EHRs).
- Unstructured clinical notes within EHRs contain valuable data for understanding complex conditions like pediatric Long COVID.
- Existing methods may not fully capture the nuances of Long COVID in children, necessitating advanced analytical approaches.
Purpose of the Study:
- To develop and validate a natural language processing (NLP) pipeline for extracting symptoms and functional impacts of Long COVID in pediatric patients from unstructured EHR data.
- To compare the prevalence of identified Long COVID concepts between pediatric patients with and without Long COVID diagnoses.
Main Methods:
- Analysis of 48,287 outpatient progress notes from 10,618 pediatric patients across 12 institutions, focusing on notes from 28 to 179 days post-COVID diagnosis.
- Development of an NLP pipeline to identify 21 symptoms and 4 functional impact categories associated with Long COVID.
- Validation of the NLP pipeline by subject matter experts (SMEs) and comparison of concept prevalence between Long COVID and acute COVID cohorts.
Main Results:
- The NLP pipeline demonstrated moderate accuracy (F1 = .80), improving to high accuracy (F1 = .90) with high-confidence SME assertions.
- The 25 identified Long COVID concept categories were significantly more prevalent in the presumptive Long COVID cohort compared to the acute COVID cohort.
- Differences were observed in concepts identified from clinical notes versus structured EHR data, highlighting the value of unstructured data.
Conclusions:
- Incorporating NLP into the analysis of unstructured EHR data provides critical insights into pediatric Long COVID syndromes.
- The developed NLP methodology is valuable for creating computable phenotypes and accurately characterizing children with Long COVID.
- This approach underscores the importance of leveraging advanced NLP techniques to enhance understanding and management of pediatric Long COVID.
Objective:
To develop a natural language processing (NLP) pipeline for unstructured electronic health record (EHR) data to identify symptoms and functional impacts associated with Long COVID in children.
Materials And Methods:
We analyzed 48 287 outpatient progress notes from 10 618 pediatric patients from 12 institutions. We evaluated notes obtained 28 to 179 days after a COVID-19 diagnosis or positive test. Two samples were examined: patients with evidence of Long COVID and patients with acute COVID but no evidence of Long COVID based on diagnostic codes. The pipeline identified clinical concepts associated with 21 symptoms and 4 functional impact categories. Subject matter experts (SMEs) screened a sample of 4586 terms from the NLP output to assess pipeline accuracy. Prevalence and concordance of each of the 25 concepts was compared between the 2 patient samples.
Results:
A binary assertion measure comparing SME and NLP assertions showed moderate accuracy (N = 4133; F1 = .80) and improved substantially when only high-confidence SME assertions were considered (N = 2043; F1 = .90). Overall, the 25 Long COVID concept categories were markedly more prevalent in the presumptive Long COVID cohort, and differences were noted between concepts identified in notes versus structured data.
Discussion:
This preliminary analysis illustrates the additional insight into a syndrome such as Long COVID gained from incorporating notes data, characterizing symptoms and functional impacts.
Conclusion:
These data support the importance of incorporating NLP methodology when possible into designing computable phenotypes and to accurately characterize patients with Long COVID.

