Related Experiment Video
Updated: Sep 11, 2025

Author Spotlight: Advancing Endoscopic Ossiculoplasty – Techniques, Innovations, and Practical Guidance for Clinical Integration
Published on: January 26, 2024
Comparing GPT-4o and o1 in Otolaryngology: An Evaluation of Guideline Adherence and Accuracy
Soumil Prasad1, Nicholas DiStefano, Nicholas Khuu
1Department of Otolaryngology, University of Miami Miller School of Medicine, Miami, FL.
Abstract:
Artificial-intelligence chatbots are gaining prominence in otolaryngology, yet their clinical safety depends on strict adherence to practice guidelines. The authors compared the accuracy of OpenAI's general-purpose GPT-4o model with the specialty-tuned o1 model on 100 otolaryngology questions drawn from national guidelines and common clinical scenarios spanning 7 subspecialty domains. Blinded otolaryngologists graded each answer as correct, partially correct, incorrect, or non-answer (scores 1, 0.5, 0, respectively), and paired statistical tests assessed performance differences. The o1 model delivered fully correct responses for 73% of questions, partially correct for 26%, and incorrect for 1%, yielding a mean accuracy score of 0.86. GPT-4o produced 64% correct and 36% partially correct answers with no incorrect responses, for a mean score of 0.82. The 4-point gap was not statistically significant (paired t test P=0.165; Wilcoxon P=0.157). Pediatric questions had the highest correctness (o1=92.9%, GPT-4o=78.6%). No domain showed systematic critical errors. Both models thus supplied predominantly guideline-concordant information, and specialty tuning conferred only a modest, nonsignificant benefit in this data set. These findings suggest contemporary large-language models may approach reliability thresholds suitable for supervised decision support in otolaryngology, but continual validation and oversight remain essential before routine deployment.
Related Concept Videos
Standards of Care II
Methods of Documentation II: POMR
Standards of Care I

