Related Experiment Video
Updated: Aug 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Utility of large language models in pediatric trauma triage: An age-stratified analysis of prompt engineering
Brendan T Fox1, Ascharya Kushidhan Balaji1, Philip Seger1
1Jacobs School of Medicine and Biomedical Sciences, University at Buffalo, Buffalo, NY, USA.
Background:
Pediatric trauma triage is complicated by age-dependent physiologic variability and low case volumes, contributing to persistent undertriage and overtriage despite standardized guidelines. Large language models (LLMs) may offer potential decision-support regardless of pediatric age group.
Methods:
We performed a retrospective analysis of 354 EMS-to-hospital pediatric trauma audio recordings from two Level I pediatric trauma centers (2023-2025). Audio recordings were transcribed and evaluated across four LLM prompting strategies (Simple Zero-Shot, L1-biased, L2-biased, Ensemble), three input formats (raw, processed, structured), and three LLMs. Performance was assessed relative to clinician triage decisions and an Injury Severity Score (ISS ≥16) reference standard. Outcomes included accuracy, sensitivity, specificity, undertriage, overtriage, and agreement (Cohen's κ).
Results:
Prompt engineering significantly improved LLM performance. Compared with Simple Zero-Shot prompting (71.5% accuracy), criteria-based strategies achieved higher accuracy, with the Ensemble approach demonstrating the most balanced performance (83.9% vs clinicians 78.8%; p = 0.026), equivalent undertriage (31.9% vs 34.0%; p = 1.0), and lower overtriage (13.7% vs 19.2%; p = 0.022). Performance remained stable across age groups (all p > 0.10). Prompt strategy exerted the largest effect on accuracy (Δ21.5%), while model selection (Δ0.7%) and input format (Δ2.2%) had minimal impact. Prompt strategies produced distinct tradeoffs between sensitivity and specificity.
Conclusions:
Structured prompt engineering enables general-domain LLMs to achieve clinician-comparable performance in pediatric trauma triage across age groups. Prompt design, rather than model selection, was the primary determinant of performance. These findings support that structured prompt engineering improves off-the-shelf LLM performance over unstructured prompting across age groups.