Related Experiment Video
Updated: May 10, 2026

05:57
A Teleoperated Robotic System-Assisted Percutaneous Transiliac-Transsacral Screw Fixation Technique
Published on: January 6, 2023
2.3K
Surgeons vs ChatGPT: Assessment and Feedback Performance Based on Real Surgical Scenarios
Cristián Jarry Trujillo1, Javier Vela Ulloa1, Gabriel Escalona Vivas1
1Experimental Surgery and Simulation Center, Department of Digestive Surgery, Pontificia Universidad Católica de Chile, Santiago, Chile.
Journal of Surgical Education
|May 15, 2024
Summary
ChatGPT demonstrates comparable error detection and feedback quality to experienced surgeons in surgical education scenarios. This artificial intelligence tool shows significant potential for improving surgical skills and training effectiveness.
Area of Science:
- Medical Education
- Artificial Intelligence in Surgery
- Surgical Skill Assessment
Background:
- Artificial intelligence (AI) is increasingly used in medical and surgical education.
- Large language models (LLMs) like ChatGPT offer potential for providing feedback to enhance surgical skills.
- This study evaluates ChatGPT's efficacy in providing feedback on surgical scenarios.
Purpose of the Study:
- To assess the ability of ChatGPT to provide constructive feedback on surgical scenarios.
- To compare ChatGPT's feedback quality and utility against that of experienced surgeons.
- To determine the potential of AI tools in surgical skill development.
Main Methods:
- Surgical scenarios were converted into neutral text narratives.
- ChatGPT 4.0 and three surgeons evaluated these texts, identifying errors and providing feedback.
- Surgical residents, an Education-Expert (EE), and a Clinical-Expert (CE) assessed the utility and quality of the feedback.
Main Results:
- ChatGPT's feedback was deemed useful in 96.43% of resident evaluations, comparable to surgeons B and C.
- ChatGPT and surgeons received similar scores for feedback quality (FQ).
- ChatGPT achieved higher FQ scores than surgeons A and B according to the EE, and comparable scores to surgeons A and C from the CE.
Conclusions:
- ChatGPT effectively identifies errors in surgical scenarios, with a detection rate similar to experienced surgeons.
- The AI-generated feedback is valuable for surgical skill improvement, performing comparably to human instructors.
- Residents, EE, and CE frequently perceived ChatGPT's feedback as human-generated, indicating its naturalistic quality.

