Related Experiment Video
Updated: May 15, 2025

08:06
Mixed Reality Assisted Radical Endoscopic Thyroidectomy
Published on: January 31, 2025
174
Artificial Intelligence Augmentation: Performance of GPT-4 and GPT-3.5 on the Plastic Surgery In-service Examination
Daniel Najafali1, Erik Reiche2, Sthefano Araya3
1From the Carle Illinois College of Medicine, University of Illinois Urbana-Champaign, Urbana, IL.
Plastic and Reconstructive Surgery. Global Open
|April 11, 2025
Summary
GPT-4 demonstrated improved performance on the Plastic Surgery In-service Examination compared to GPT-3.5, but still lags behind experienced residents. Further development is needed for AI to be a reliable surgical education tool.
Area of Science:
- Artificial Intelligence in Medical Education
- Large Language Models in Surgery
- Plastic Surgery Training
Background:
- ChatGPT-3.5 achieved a 52nd percentile score on the Plastic Surgery In-service Examination, equivalent to a first-year resident.
- The advanced GPT-4 model, with its larger training dataset, was evaluated for potentially enhanced performance.
- This study hypothesized that GPT-4 would surpass GPT-3.5, offering greater utility in surgical education.
Purpose of the Study:
- To assess the performance of GPT-4 and GPT-3.5 on the 2022 Plastic Surgery In-service Examination.
- To compare AI performance against national metrics for plastic surgery residents.
- To determine the potential of GPT-4 as a tool for surgical education.
Main Methods:
- GPT-4 and GPT-3.5 were administered questions from the 2022 Plastic Surgery In-service Examination.
- Three distinct prompting strategies were employed for both AI models.
- Performance was benchmarked against the 2022 American Society of Plastic Surgeons Norm Tables.
Main Results:
- GPT-4 achieved an overall accuracy of 63%, outperforming GPT-3.5's 58% accuracy.
- The highest accuracy for GPT-4 was 68% using multiple-choice questions with explanations, achieving the 93rd percentile for first-year residents.
- Despite improvements, GPT-4's best performance (68%) placed it at the 15th percentile for sixth-year residents.
Conclusions:
- GPT-4 significantly outperformed GPT-3.5 on the Plastic Surgery In-service Examination.
- Current AI performance, while improved, does not yet match that of experienced residents or attending surgeons.
- Further AI model refinement is necessary to establish its value as a comprehensive surgical education resource.

