Related Experiment Video
Updated: Apr 8, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Comparing Large Language Models' Performances on Otolaryngology Knowledge Assessment Questions
Ryan Cook1, Abner Kahan1, Thomas Scharfenberger1
1Albert Einstein College of Medicine, Bronx, New York, United States.
Applied Clinical Informatics
|April 6, 2026
Summary
This study tested large language models (LLMs) on otolaryngology knowledge. Top models achieved around 76% accuracy, indicating a plateau for general LLMs in specialized medical fields.
Area of Science:
- Medical Education
- Artificial Intelligence
- Otolaryngology
Background:
- Large language models (LLMs) show potential in medical education.
- Evaluating LLM performance on specialized medical knowledge is crucial.
- Otolaryngology knowledge assessment requires accurate AI tools.
Purpose of the Study:
- To assess the performance of OpenAI's GPT-4 Turbo and 10 other commercial large language models (LLMs) on specialized otolaryngology knowledge.
- To compare the utility of these LLMs in otolaryngology medical education.
- To identify the current capabilities and limitations of general-purpose LLMs in a specific medical domain.
Main Methods:
- 1,075 otolaryngology questions from OTO QUEST were administered to GPT-4 Turbo using a zero-shot approach.
- Accuracy was analyzed using logistic regression, controlling for question difficulty, year, and subspecialty.
- Comparative analysis involved 10 commercial LLMs, including Claude-3.5-Sonnet, Gemini-1.5-Pro, and GPT-4o, using Cochran's Q test and McNemar's pairwise comparison.
Main Results:
- GPT-4 Turbo achieved 72.09% accuracy, excelling in Practice Management but declining with moderate and hard difficulty questions.
- In comparative analysis, Grok-3 (76.3%), Claude-3.5-Sonnet (73.0%), and GPT-4o (69.9%) showed higher accuracy than GPT-4 Turbo.
- The top-performing models demonstrated an accuracy plateau around 73-76% on this specialized medical knowledge dataset.
Conclusions:
- Current general-purpose LLMs demonstrate promising but limited capabilities in assessing specialized otolaryngology knowledge.
- An accuracy plateau exists for these models, suggesting a need for domain-specific training.
- Further research into specialized training for LLMs is recommended to improve performance in medical education and practice.

