A natural language processing pipeline for identifying pediatric long COVID symptoms and functional impacts in

H Timothy Bunnell1, Cara Reedy1, Vitaly Lorman2

  • 1Biomedical Research Informatics Center, Nemours Children's Health, Wilmington, DE 19803, United States.

JAMIA Open
|September 8, 2025
PubMed

Insights

A new natural language processing (NLP) pipeline accurately identifies Long COVID symptoms and functional impacts in children using electronic health records. This method enhances understanding of pediatric Long COVID, improving patient characterization.

Area of Science:

  • Pediatric Health
  • Medical Informatics
  • Natural Language Processing

Background:

  • Long COVID presents a significant challenge in pediatric populations, with symptoms and functional impacts often poorly characterized in electronic health records (EHRs).
  • Unstructured clinical notes within EHRs contain valuable data for understanding complex conditions like pediatric Long COVID.
  • Existing methods may not fully capture the nuances of Long COVID in children, necessitating advanced analytical approaches.

Purpose of the Study:

  • To develop and validate a natural language processing (NLP) pipeline for extracting symptoms and functional impacts of Long COVID in pediatric patients from unstructured EHR data.
  • To compare the prevalence of identified Long COVID concepts between pediatric patients with and without Long COVID diagnoses.

Main Methods:

  • Analysis of 48,287 outpatient progress notes from 10,618 pediatric patients across 12 institutions, focusing on notes from 28 to 179 days post-COVID diagnosis.
  • Development of an NLP pipeline to identify 21 symptoms and 4 functional impact categories associated with Long COVID.
  • Validation of the NLP pipeline by subject matter experts (SMEs) and comparison of concept prevalence between Long COVID and acute COVID cohorts.

Main Results:

  • The NLP pipeline demonstrated moderate accuracy (F1 = .80), improving to high accuracy (F1 = .90) with high-confidence SME assertions.
  • The 25 identified Long COVID concept categories were significantly more prevalent in the presumptive Long COVID cohort compared to the acute COVID cohort.
  • Differences were observed in concepts identified from clinical notes versus structured EHR data, highlighting the value of unstructured data.

Conclusions:

  • Incorporating NLP into the analysis of unstructured EHR data provides critical insights into pediatric Long COVID syndromes.
  • The developed NLP methodology is valuable for creating computable phenotypes and accurately characterizing children with Long COVID.
  • This approach underscores the importance of leveraging advanced NLP techniques to enhance understanding and management of pediatric Long COVID.
Abstract