Related Experiment Video
Updated: Jun 25, 2025

08:08
Simulator Training for Endovascular Neurosurgery
Published on: May 6, 2020
3.6K
Performance of ChatGPT on American Board of Surgery In-Training Examination Preparation Questions
Catherine G Tran1, Jeremy Chang1, Scott K Sherman1
1Department of Surgery, University of Iowa Hospitals & Clinics, Iowa City, Iowa.
The Journal of Surgical Research
|May 24, 2024
Summary
ChatGPT achieved 62% accuracy on Surgical Council on Resident Education (SCORE) self-assessment questions. While proficient in recall, it struggled with complex clinical decision-making, necessitating caution with its responses in general surgery education.
Area of Science:
- Artificial Intelligence in Medical Education
- Natural Language Processing in Healthcare
- Surgical Training and Assessment
Background:
- Large language models (LLMs) like Chat Generative Pretrained Transformer (ChatGPT) demonstrate human-like text generation capabilities.
- The integration of AI tools in medical education necessitates rigorous performance evaluation.
- Assessing ChatGPT's utility for resident self-assessment in surgery is crucial for understanding its potential role.
Purpose of the Study:
- To evaluate the performance of ChatGPT (GPT-3.5) on Surgical Council on Resident Education (SCORE) self-assessment multiple-choice questions.
- To identify specific surgical topics where ChatGPT exhibits strengths and weaknesses.
- To analyze the quality and accuracy of ChatGPT's explanations for its answers.
Main Methods:
- Random selection of general surgery multiple-choice questions from the SCORE question bank.
- ChatGPT (GPT-3.5, April-May 2023) was utilized to answer the selected questions.
- Accuracy of ChatGPT's responses was recorded and analyzed across different surgical subspecialties.
Main Results:
- ChatGPT answered 123 out of 200 questions correctly, achieving an overall accuracy of 62%.
- Performance varied significantly by topic, with lower scores in biliary (25%), surgical critical care (30%), and higher scores in biostatistics (100%) and fluid/electrolytes/acid-base (100%).
- Incorrect answers often featured plausible but factually inaccurate information regarding anatomy and surgical steps; two questions were unanswerable due to insufficient data.
Conclusions:
- ChatGPT demonstrates a 62% accuracy rate on SCORE self-assessment questions, indicating potential but limited utility.
- The model excels at recall-based questions but falters in complex clinical decision-making scenarios.
- Caution is advised when using ChatGPT for general surgery self-assessment due to the risk of confidently presented, yet inaccurate information.

