Related Experiment Video
Updated: Sep 4, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Real-World Evaluation of Large Language Model-Generated Hospital Courses in Pediatrics
Jasmine E Kim1,2, Jonathan D Hron2,3, Daniel J Kats2,3
1Department of Psychiatry, Boston Children's Hospital, Boston, Massachusetts.
Background:
Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized.
Objective:
To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and LLM-generated hospital courses in pediatric inpatient care during early clinical implementation.
Methods:
We conducted a descriptive evaluation from June 10 to August 8, 2025, at an academic freestanding children's hospital using an Epic EHR with an integrated LLM tool (GPT-4o and GPT-4.1). Clinicians across multiple roles, including attending physicians, residents, and advanced practice providers, reviewed LLM-generated hospital courses for their own patients. Clinicians identified and categorized errors (hallucinations, inaccuracies, or omissions). They also rated text quality (comprehensiveness, conciseness, coherence) on a 5-point scale and perceived harm on an 8-point scale.
Results:
A total of 129 LLM-generated hospital courses were reviewed (median length of stay, 3 days; IQR, 2-7) by 50 involved clinicians. Hallucinations occurred in 21% (95% CI, 14%-29%) of the hospital courses, inaccuracies in 41% (53/129; 95% CI, 33%-50%), and omissions in 24% (31/129; 95% CI, 17%-32%). Overall, perceived harm ratings were low (median, 0; IQR, 0-1). Text quality ratings were high (median [IQR]: comprehensiveness, 4 [3-5]; conciseness, 4 [4-5]; coherence, 4 [4-5]) and comparable with prior literature.
Conclusion:
In this pediatric evaluation of LLM-generated hospital courses reviewed by frontline clinicians, errors were common, but perceived potential harm was low, even assuming use without clinician correction. These findings support the use of LLM-generated hospital courses as starting drafts when paired with clinician review and institutional safeguards.
