Related Experiment Video
Updated: May 31, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
When algorithms triage trauma: Diagnostic accuracy, undertriage risk, and prompt fragility in frontier large language
Yahya Kemal Çalışkan1, Fatih Başak2, Olgun Erdem2
1Department of General Surgery, Kanuni Training and Research Hospital, University of Health Sciences, Istanbul, Turkey.
Background:
Large language models are increasingly proposed for emergency decision support, yet high aggregate accuracy may conceal clinically meaningful subgroup errors. This study specifically focused on undertriage because trauma triage is vulnerable to false reassurance, and small misclassifications may delay escalation, imaging, transfer, and trauma-team activation. We hypothesized that strong overall discrimination would coexist with clinically relevant subgroup vulnerability and prompt fragility.
Methods:
This observational in silico benchmarking study, meaning a computer-based comparison performed on standardized written cases rather than live patients, used 150 trauma vignettes derived from deidentified, publicly available educational and registry-informed materials. The design primarily emulated secondary triage in a referring emergency department setting. Three frontier large language model families were assessed between February 4 and February 18, 2026, using a locked prompt template: GPT-5.2 (OpenAI), Gemini 3.0 Flash (Google), and Claude 4.5 Opus (Anthropic). The reference standard was consensus adjudication by board-certified surgeons using Injury Severity Score-anchored categories supplemented by urgent management priority. The primary outcome was binary correct identification of major trauma versus nonmajor trauma; secondary outcomes included sensitivity, specificity, area under the receiver operating characteristic curve, Cohen's kappa, subgroup undertriage, prompt perturbation sensitivity, and simulated downstream delay.
Results:
Overall sensitivity for major trauma was 89.4%, specificity was 86.5%, and the pooled kappa was 0.74. GPT-5.2 achieved the highest discrimination with an area under the receiver operating characteristic curve of 0.94 and a kappa of 0.82; Gemini 3.0 Flash achieved an area under the receiver operating characteristic curve of 0.91 and a kappa of 0.75; and Claude 4.5 Opus achieved an area under the receiver operating characteristic curve of 0.89 and a kappa of 0.65. Undertriage occurred disproportionately in geriatric scenarios (17.6% vs 5.2% in younger adults; P = .012) and was more frequent in blunt trauma than in penetrating trauma (11.8% vs 4.1%; P = .045). Across standardized lexical perturbations applied to all 150 vignettes, 11.2% of repeated prompts changed triage category, with a maximum instability of 24.5%. Simulated delay linked to undertriage affected 14.2% of major-trauma cases and yielded a mean theoretical delay of 18.5 minutes.
Conclusion:
Frontier large language models showed strong overall discrimination but clinically relevant fragility. Their errors were not random; they clustered in older adults and less overt trauma patterns, the very scenarios in which real-world trauma systems already struggle. These findings support large language model use, if at all, as supervised adjuncts rather than autonomous triage agents.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy