Related Experiment Video
Updated: Apr 9, 2026

Utilizing a 3D Printed Laparoscopic Nissen Fundoplication Model to Shorten a Resident's Learning Curve
Published on: August 15, 2025
Large Language Models Outperform PGY-5 Residents on the Orthopaedic In-Training Examination: A Comparative Analysis
Rushil Dave1, Nimit Vediya, Naman Sharma
1From the Temerty Faculty of Medicine (Dave, Vediya, Sharma), University of Toronto, Toronto, ON, Canada, the Division of Orthopaedic Surgery (Shah, Whelan, Wolfstadt), Department of Surgery, University of Toronto, Toronto, ON, Canada, the Division of Orthopaedic Surgery (Whelan), Department of Surgery, St. Michael's Hospital, Unity Health Network, Toronto, ON, Canada, the University of Toronto Orthopaedic Sports Medicine (Whelan), Toronto, ON, Canada, and the Division of Orthopaedic Surgery (Wolfstadt), Department of Surgery, Mount Sinai Hospital, Toronto, ON, Canada.
Introduction:
Large language models (LLMs), such as ChatGPT, are becoming increasingly prevalent, particularly in medical education and clinical assessments. Previous LLMs were seen to perform at the level of a first-year resident on the 2022 Orthopaedic In-Training Examination (OITE). With exponential advances in LLMs over the past 3 years, the true capabilities of these models remain unexplored. In addition, the addition of image processing further increases their clinical applicability. The purpose of this study was to evaluate the performance of six LLMs on the 2024 OITE.
Methods:
Six LLMs were evaluated in this study: ChatGPT (GPT-4o), Gemini 2.0 Flash, Grok 3, Mistral Large 2.7, DeepSeek R1, and Llama. ChatGPT, Gemini, Grok, and Mistral could evaluate images and text while DeepSeek and Llama were limited to text. Accuracy, image interpretation, and logical consistency were assessed in 203 multiple-choice questions, stratified by difficulty and type of the question. Statistical analyses involved chi-square tests, Fisher exact tests, z-tests, and Cohen κ tests.
Results:
ChatGPT performed with the highest accuracy (74.9%), followed by DeepSeek, Llama, Grok, Mistral, and Gemini. ChatGPT also led in logical consistency (72.4%) and image interpretation (73.8%). Logical consistency strongly correlated with accuracy and correctness ( P < 0.00001). As difficulty increased, performance declined across all models.
Conclusion:
ChatGPT consistently scored the highest in terms of accuracy across all metrics while also maintaining reasoning quality. Compared with resident averages, ChatGPT performed at a postgraduate year five level which indicates its potential for integration into orthopaedic clinics, electronic medical records, and surgical planning. Further development models would allow for better performance on difficult questions and creating orthopaedic focused models could enhance these results.

