Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing acuity in pediatric emergency department triage: performance of a large language model
Kush Narang1, Newton Addo1, Christopher Y K Williams2,3
1Department of Emergency Medicine, University of California, San Francisco, San Francisco, CA, USA.
Insights
Large language models (LLMs) show moderate accuracy in identifying higher-acuity pediatric patients from clinical notes. However, LLMs may prioritize younger children, requiring pediatric-specific evaluation before clinical use.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Pediatric Emergency Medicine
Background:
- Pediatric triage accuracy varies significantly across emergency departments (EDs).
- Consistent decision-making in pediatric emergency care remains a challenge.
- Large language models (LLMs) are being explored to standardize triage processes.
Purpose of the Study:
- To evaluate the performance of a specific LLM (GPT-5-mini) in identifying higher-acuity children from clinical notes.
- To assess the LLM's accuracy in a large cohort of pediatric emergency department visits.
Main Methods:
- De-identified clinical notes from 228,104 pediatric ED visits were used.
- An LLM (GPT-5-mini) was tasked with identifying the higher-acuity child from paired notes.
- Statistical analysis was performed to determine accuracy and identify factors influencing performance.
Main Results:
- The LLM achieved an overall accuracy of 0.73 in identifying higher-acuity children.
- Accuracy improved with greater differences in acuity between paired visits.
- The LLM was less accurate when the higher-acuity child was older or when age differences were large.
Conclusions:
- LLMs demonstrate moderate accuracy for pediatric acuity assessment but exhibit biases similar to human triage.
- The tendency to prioritize younger children warrants further investigation.
- Pediatric-specific LLM evaluation and optimization are crucial before clinical implementation.
Abstract:
Pediatric triage performance varies across emergency departments (ED), contributing to ongoing challenges in pediatric emergency care. There is growing interest in using large language models (LLMs) to support more consistent triage decision-making in children. We evaluated an LLM's (GPT-5-mini) ability to identify the higher-acuity child from pairs of de-identified clinical notes. Across 228,104 pediatric ED visits, the LLM achieved an overall accuracy of 0.73 (95% CI, 0.73-0.74) in identifying the higher-acuity child, with accuracy improving as acuity differences between visits increased. The LLM was less likely to be correct when the higher-acuity child was older (odds ratio, 0.62, 95% CI, 0.61-0.63) and when the age difference between children was large (0.75, 95% CI, 0.70-0.79). The LLM showed moderate overall accuracy in assessing pediatric acuity and demonstrated a tendency to prioritize younger children, similar to human performance. These findings highlight the need for pediatric-specific LLM evaluation and optimization before clinical use.
