Evaluating Large Language Models for Transparent Quality-of-Care Measurement in Children with ADHD
Yair Bannett1, Malvika Pillai2, Tracy X Huang1
1Division of Developmental-Behavioral Pediatrics, Stanford University School of Medicine, Stanford, California, USA.
Insights
Large language models (LLMs) accurately identify parent training in behavior management (PTBM) recommendations in ADHD clinical notes. This offers a scalable method for quality measurement, surpassing manual chart review limitations.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Informatics
- Pediatric Medicine
Background:
- Guideline-concordant care for pediatric ADHD emphasizes parent training in behavior management (PTBM) as a first-line treatment.
- Manual chart review for assessing guideline adherence is resource-intensive, hindering scalable quality measurement.
Purpose of the Study:
- To assess the accuracy and explainability of large language models (LLMs) in identifying PTBM recommendations within pediatric electronic health record (EHR) notes.
- To establish LLMs as a viable, scalable alternative to manual chart review for quality-of-care assessment in pediatric ADHD.
Main Methods:
- A retrospective cohort study analyzed clinical notes from children aged 4-6 with ADHD diagnoses across 27 primary care clinics.
- Three generative LLMs (Claude-3.5, GPT-4o, LLaMA-3.3-70B) evaluated assessment and plan sections for PTBM recommendations.
- Model performance was measured by sensitivity, PPV, and F1-score, with explainability assessed using the QUEST framework.
Main Results:
- All evaluated LLMs demonstrated high accuracy in identifying PTBM recommendations, comparable to expert chart review.
- Claude-3.5 achieved the best balance of performance (sensitivity=0.89, PPV=0.95, F1=0.92) and explainability.
- LLMs identified that 26.4% of young ADHD patients had documented PTBM recommendations at their initial visit.
Conclusions:
- LLMs can reliably extract guideline-concordant ADHD treatment recommendations from unstructured EHR data.
- Incorporating LLM explainability is crucial for transparent and scalable quality measurement in healthcare.
- This technology offers a promising solution for improving the monitoring of non-pharmacological ADHD treatment adherence.
Importance:
Guideline-concordant care for young children with attention-deficit/hyperactivity disorder (ADHD) includes recommending parent training in behavior management (PTBM) as first-line treatment. However, assessing guideline adherence through manual chart review is time-consuming and costly, limiting scalable and timely quality-of-care measurement.
Objective:
To evaluate the accuracy and explainability of large language models (LLMs) in identifying PTBM recommendations in pediatric electronic health record (EHR) notes as a scalable alternative to manual chart review.
Design Setting And Participants:
This retrospective cohort study was conducted in a community-based pediatric healthcare network in California consisting of 27 primary care clinics. The study cohort included children aged 4-6 years with ≥ 2 primary care visits between 2020-2024 and ICD-10 diagnoses of ADHD or ADHD symptoms (n=542 patients). Clinical notes from the first ADHD-related visit were included. A stratified subset of 122 notes, including all cases with model disagreement, was manually annotated to assess model performance in identifying PTBM recommendations and rank model explanations.
Exposures:
Assessment and plan sections of clinical notes were analyzed using three generative large language models (Claude-3.5, GPT-4o, and LLaMA-3.3-70B) to identify the presence of PTBM recommendations and generate explanatory rationales and documentation evidence.
Main Outcomes And Measures:
Model performance in identifying PTBM recommendations (measured by sensitivity, positive predictive value (PPV), and F1-score) and qualitative explainability ratings of model-generated rationales (based on the QUEST framework).
Results:
All three models demonstrated high performance compared to expert chart review. Claude-3.5 showed balanced performance (sensitivity=0.89, PPV=0.95, and F1-score=0.92) and ranked highest in explainability. LLaMA3.3-70B achieved sensitivity=0.91, PPV=0.89, and F1-score=0.90, ranking second for explainability. GPT-4o had the highest PPV [0.97] but lowest sensitivity [0.82], with an F1-score of 0.89 and the lowest explainability ranking. Based on classifications from the best-performing model, Claude-3.5, 26.4% (143/542) of patients had documented PTBM recommendations at their first ADHD-related visit.
Conclusions And Relevance:
LLMs can accurately extract guideline-concordant clinician recommendations for non-pharmacological ADHD treatment from unstructured clinical notes while providing clear explanations and supporting evidence. Evaluating model explainability as part of LLM implementation for medical chart review tasks can promote transparent and scalable solutions for quality-of-care measurement.
Related Concept Videos
Attention-Deficit/Hyperactivity Disorder
Diagnostic Criteria and Symptoms
To diagnose ADHD, symptoms must manifest before age 12 and be evident across multiple settings....
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...


