Related Experiment Video
Updated: May 31, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
484
Advancements in AI Medical Education: Assessing ChatGPT's Performance on USMLE-Style Questions Across Topics and
Parker Penny1, Riley Bane2, Valerie Riddle1
1Medical Education, University of South Florida Morsani College of Medicine, Tampa, USA.
Cureus
|January 24, 2025
Summary
ChatGPT-4 significantly outperformed ChatGPT-3.5 on USMLE-style questions, surpassing human performance. This artificial intelligence model shows promise for medical education, demonstrating improved accuracy and consistency across various medical topics and question difficulties.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
Background:
- Artificial intelligence (AI) language models demonstrate variable performance on medical licensing exams.
- Previous AI models have passed some imageless USMLE tests but failed specialty-specific exams, suggesting topic or difficulty-based limitations.
Purpose of the Study:
- To evaluate the performance of ChatGPT-3.5 and ChatGPT-4 on USMLE-style questions across diverse medical topics.
- To compare AI model performance against human test-takers (AMBOSS users) regarding medical topic and question difficulty.
Main Methods:
- 900 USMLE-style multiple-choice questions from AMBOSS, excluding image/chart/table-based items, were used.
- Questions were divided into 18 topics and categorized by exam type (Step 1 vs. Step 2).
- ChatGPT-3.5 and ChatGPT-4 were tested over multiple trials, with performance compared to human users.
Main Results:
- ChatGPT-4 achieved 71.33% accuracy, surpassing AMBOSS users (54.38%) and ChatGPT-3.5 (46.23%).
- ChatGPT-4 demonstrated significantly greater accuracy and concordance between trials compared to ChatGPT-3.5 (p<.001).
- Both AI models and human users showed decreased performance with increased question difficulty, but ChatGPT-4's accuracy declined less significantly.
Conclusions:
- ChatGPT-4 represents a substantial advancement over ChatGPT-3.5 for medical education applications.
- ChatGPT-4 surpassed human performance on the tested USMLE-style questions, indicating its potential as a valuable learning tool.
- AI performance variability appears linked to question complexity rather than specific medical topic knowledge gaps.

