Related Experiment Video
Updated: Jun 22, 2025

03:40
Acupoint Catgut Embedding Therapy in Traditional Chinese Medicine for Managing Allergic Rhinitis
Published on: December 20, 2024
439
Comparative Performance of ChatGPT 3.5 and GPT4 on Rhinology Standardized Board Examination Questions.
Evan A Patel1, Lindsay Fleischer1, Peter Filip1
1Department of Otorhinolaryngology-Head and Neck Surgery Rush University Medical Center Chicago Illinois USA.
OTO Open
|June 28, 2024
Summary
Large language models (LLMs) like ChatGPT 3.5 and GPT4 were tested on Otolaryngology board exam questions. GPT4 performed significantly better than ChatGPT 3.5, suggesting AI's growing potential in medical education.
Area of Science:
- Artificial Intelligence in Medicine
- Deep Learning Applications
- Medical Education Technology
Background:
- Large language models (LLMs) such as ChatGPT have emerged due to advances in artificial intelligence (AI).
- Evaluating the performance of these AI models on specialized medical examinations is crucial for understanding their potential applications.
Purpose of the Study:
- To assess the performance of ChatGPT 3.5 and GPT4 on Otolaryngology (Rhinology) Standardized Board Examination questions.
- To compare AI performance against Otolaryngology residents' results.
Main Methods:
- 127 rhinology standardized questions from BoardVitals were used.
- ChatGPT 3.5 and GPT4 answered 93 text-based questions; GPT4 also answered 34 image-based questions.
- Performance was compared to resident performance, with a pass-fail cutoff at the 10th percentile.
Main Results:
- ChatGPT 3.5 achieved 45.2% accuracy (8th percentile) on text questions, likely failing the exam.
- GPT4 achieved 86.0% accuracy (66th percentile) on text questions and 64.7% on image questions, indicating a strong chance of passing.
- Statistical significance was found for both models' performance (P=.0001 for 3.5, P=.001 for 4).
Conclusions:
- ChatGPT 3.5 is unlikely to pass the American Board of Otolaryngology Written Question Exam (ABOto WQE).
- GPT4 demonstrates a significantly higher likelihood of passing the ABOto WQE.
- The rapid advancement of AI suggests a potential future role in otolaryngology education and practice.

