Evaluating Large Language Models for Transparent Quality-of-Care Measurement in Children with ADHD

Yair Bannett1, Malvika Pillai2, Tracy X Huang1

  • 1Division of Developmental-Behavioral Pediatrics, Stanford University School of Medicine, Stanford, California, USA.

Insights

Large language models (LLMs) accurately identify parent training in behavior management (PTBM) recommendations in ADHD clinical notes. This offers a scalable method for quality measurement, surpassing manual chart review limitations.

Area of Science:

  • Artificial Intelligence in Healthcare
  • Clinical Informatics
  • Pediatric Medicine

Background:

  • Guideline-concordant care for pediatric ADHD emphasizes parent training in behavior management (PTBM) as a first-line treatment.
  • Manual chart review for assessing guideline adherence is resource-intensive, hindering scalable quality measurement.

Purpose of the Study:

  • To assess the accuracy and explainability of large language models (LLMs) in identifying PTBM recommendations within pediatric electronic health record (EHR) notes.
  • To establish LLMs as a viable, scalable alternative to manual chart review for quality-of-care assessment in pediatric ADHD.

Main Methods:

  • A retrospective cohort study analyzed clinical notes from children aged 4-6 with ADHD diagnoses across 27 primary care clinics.
  • Three generative LLMs (Claude-3.5, GPT-4o, LLaMA-3.3-70B) evaluated assessment and plan sections for PTBM recommendations.
  • Model performance was measured by sensitivity, PPV, and F1-score, with explainability assessed using the QUEST framework.

Main Results:

  • All evaluated LLMs demonstrated high accuracy in identifying PTBM recommendations, comparable to expert chart review.
  • Claude-3.5 achieved the best balance of performance (sensitivity=0.89, PPV=0.95, F1=0.92) and explainability.
  • LLMs identified that 26.4% of young ADHD patients had documented PTBM recommendations at their initial visit.

Conclusions:

  • LLMs can reliably extract guideline-concordant ADHD treatment recommendations from unstructured EHR data.
  • Incorporating LLM explainability is crucial for transparent and scalable quality measurement in healthcare.
  • This technology offers a promising solution for improving the monitoring of non-pharmacological ADHD treatment adherence.
Abstract