Related Experiment Video
Updated: Sep 16, 2025

07:48
Eye Tracking During A Complex Aviation Task For Insights Into Information Processing
Published on: April 4, 2025
578
GPT-4 as a Board-Certified Surgeon: A Pilot Study
Joshua A Roshal1,2, Caitlin Silvestri3, Tejas Sathe3
1University of Texas Medical Branch, 301 University Blvd, Galveston, TX 77551 USA.
Medical Science Educator
|July 8, 2025
Summary
Large language models like GPT-4 show promise for surgical education, excelling in multiple-choice tests but struggling with complex oral board scenarios. Further development is needed for clinical decision-making applications.
Area of Science:
- Artificial Intelligence in Medical Education
- Surgical Training Technologies
- Large Language Models (LLMs)
Background:
- Large language models (LLMs) show potential for revolutionizing surgical education.
- Skepticism regarding LLM accuracy and reliability hinders adoption in medical training.
- GPT-4's performance on multiple-choice questions is known, but its clinical judgment in oral examinations is less understood.
Purpose of the Study:
- To evaluate GPT-4's general surgery knowledge using mock written and oral board-style examinations.
- To identify areas for improvement in LLMs for surgical education and practice.
- To assess the reliability of GPT-4 in simulating high-stakes clinical decision-making scenarios.
Main Methods:
- GPT-4 answered 250 multiple-choice questions (MCQs) from the Surgical Council on Resident Education (SCORE) question bank.
- GPT-4 navigated 4 oral board scenarios based on Entrustable Professional Activities (EPA) topic list.
- Responses were independently assessed for accuracy by two former oral board examiners.
Main Results:
- GPT-4 achieved 78.8% accuracy on MCQs, indicating a 92% probability of passing the American Board of Surgery Qualifying Examination (ABS QE).
- GPT-4 committed critical failures in 75% of oral board scenarios (3 out of 4 cases).
- Key failure points included incorrect timing of interventions and inappropriate surgical recommendations.
Conclusions:
- GPT-4's high MCQ performance aligns with previous findings, but it demonstrated limitations in generating accurate long-form content for oral examinations.
- Significant improvements are required in LLM performance for complex clinical decision-making.
- Future research should focus on specialized datasets and advanced reinforcement learning to enhance LLM capabilities in surgical contexts.

