Related Experiment Video
Updated: Jun 23, 2025

Author Spotlight: Enhancing Upper Limb Rehabilitation in Stroke Patients Through Advanced Robotic and Neuromodulation Technologies
Published on: October 11, 2024
Comparison of Artificial Intelligence to Resident Performance on Upper-Extremity Orthopaedic In-Training Examination
Yagiz Ozdag1, Daniel S Hayes1, Gabriel S Makar1
1Department of Orthopaedic Surgery, Geisinger Musculoskeletal Institute, Geisinger Commonwealth School of Medicine, Danville, PA.
Purpose:
Currently, there is a paucity of prior investigations and studies examining applications for artificial intelligence (AI) in upper-extremity (UE) surgical education. The purpose of this investigation was to assess the performance of a novel AI tool (ChatGPT) on UE questions on the Orthopaedic In-Training Examination (OITE). We aimed to compare the performance of ChatGPT to the examination performance of hand surgery residents.
Methods:
We selected questions from the 2020-2022 OITEs that focused on both the hand and UE as well as the shoulder and elbow content domains. These questions were divided into two categories: those with text-only prompts (text-only questions) and those that included supplementary images or videos (media questions). Two authors (B.K.F. and G.S.M.) converted the accompanying media into text-based descriptions. Included questions were inputted into ChatGPT (version 3.5) to generate responses. Each OITE question was entered into ChatGPT three times: (1) open-ended response, which requested a free-text response; (2) multiple-choice responses without asking for justification; and (3) multiple-choice response with justification. We referred to the OITE scoring guide for each year in order to compare the percentage of correct AI responses to correct resident responses.
Results:
A total of 102 UE OITE questions were included; 59 were text-only questions, and 43 were media-based. ChatGPT correctly answered 46 (45%) of 102 questions using the Multiple Choice No Justification prompt requirement (42% for text-based and 44% for media questions). Compared to ChatGPT, postgraduate year 1 orthopaedic residents achieved an average score of 51% correct. Postgraduate year 5 residents answered 76% of the same questions correctly.
Conclusions:
ChatGPT answered fewer UE OITE questions correctly compared to hand surgery residents of all training levels.
Clinical Relevance:
Further development of novel AI tools may be necessary if this technology is going to have a role in UE education.
More Related Videos
05:52Mobile Game-based Virtual Reality Program for Upper Extremity Stroke Rehabilitation
Published on: March 8, 2018
04:49Author Spotlight: Enhancing Post-Stroke Upper Limb Rehabilitation with Robotic Technologies for Improved Motor Recovery and Functional Outcomes
Published on: September 6, 2024