Related Experiment Video
Updated: Jul 9, 2025

06:04
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
409
Assessment of Artificial Intelligence Performance on the Otolaryngology Residency In-Service Exam
Arushi P Mahajan1, Christina L Shabet1, Joshua Smith2
1University of Michigan Medical School Ann Arbor Michigan USA.
OTO Open
|November 30, 2023
Summary
Large language models show limited reliability for otolaryngology medical education. ChatGPT answered only 53% of practice questions correctly, highlighting current AI limitations for surgical trainees.
Area of Science:
- Medical Education
- Artificial Intelligence in Medicine
- Otolaryngology
Background:
- Sub-specialized medical fields require accurate and comprehensive learning resources.
- Surgical trainees and learners need reliable tools for exam preparation.
- The efficacy of artificial intelligence (AI) in medical education is under investigation.
Purpose of the Study:
- To evaluate the reliability and potential use of a large language model (LLM) for answering otolaryngology-head and neck surgery practice questions.
- To assess the current efficacy of AI in providing accurate answers and explanations for surgical learners.
Main Methods:
- A public, paid-access otolaryngology question bank was utilized.
- Questions were manually inputted into ChatGPT.
- ChatGPT outputs were compared against the question bank's benchmark answers and explanations for accuracy and comprehensiveness.
Main Results:
- ChatGPT achieved a correct answer rate of 53% and a correct explanation rate of 54%.
- Accuracy of both answers and explanations decreased as question difficulty increased.
- The LLM demonstrated significant limitations in providing reliable medical education content.
Conclusions:
- Current AI-driven learning platforms are not sufficiently robust for reliable medical education in sub-specialty areas.
- AI tools may not be dependable for assisting learners in complex, sub-specialty-specific patient decision-making scenarios.
- Further development is needed to enhance AI accuracy and comprehensiveness for medical training.

