Related Experiment Video
Updated: Jun 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The Utility of Large Language Models to Assist With Emergency Triage Decisions Within Otolaryngology
Sholem Hack1, Rebecca Attal1, Rachel J Steckbeck2
1City St. George's University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center, Ramat Gan, Israel.
Objective:
To determine whether contemporary large language models can match clinician performance in evaluating the urgency of emergency otolaryngology referrals.
Study Design:
Blinded cross-sectional diagnostic reasoning study.
Setting:
Simulated emergency referral environment modeled on tertiary care otolaryngology practice.
Methods:
Thirty emergency referral scenarios spanning the spectrum of otolaryngologic urgency were independently evaluated by 4 large language models (GPT-5, GPT-4, DeepSeek, and Grok) and 4 clinicians (otolaryngology attending and resident, emergency attending and resident). Outputs were anonymized and scored by 10 blinded otolaryngologists for appropriateness of urgency and quality of explanation using a three-point scale. Statistical analyses included nonparametric group comparisons, adjusted ordinary least squares modeling with case-level control, and correlation of each entity's case profile with that of the otolaryngology attending.
Results:
Inter-rater reliability was excellent. The otolaryngology attending achieved the highest overall performance. GPT-5 demonstrated comparable mean performance, with no statistically significant difference in either domain. GPT-4 scored modestly lower but received higher mean ratings than both emergency clinicians. DeepSeek and the otolaryngology resident demonstrated intermediate performance, while Grok and the emergency clinicians performed lowest. Group-level analyses showed no significant difference between the large language model and otolaryngology cohorts; both were rated higher than emergency clinicians in this sample.
Conclusion:
GPT-5 demonstrated triage performance comparable to the otolaryngology attending in this controlled sample. Large language models may support emergency decision-making and education when specialist consultation is limited, but require supervision, transparency, and local calibration.
Related Concept Videos
Cardiopulmonary Resuscitation II: ACLS Airway Management
Cardiopulmonary Resuscitation V: Advanced Airway Management Techniques