Related Experiment Video
Updated: Jun 13, 2025

07:24
Autologous Microfractured and Purified Adipose Tissue for Arthroscopic Management of Osteochondral Lesions of the Talus
Published on: January 23, 2018
10.3K
Artificial Intelligence in Orthopaedics: Performance of ChatGPT on Text and Image Questions on a Complete AAOS
Daniel S Hayes1, Brian K Foster1, Gabriel Makar1
1Department of Orthopaedic Surgery, Geisinger Commonwealth School of Medicine, Geisinger Musculoskeletal Institute, Danville, PA.
Journal of Surgical Education
|September 16, 2024
Summary
Artificial intelligence (AI) chatbots like ChatGPT performed poorly on the 2019 Orthopaedic In-Training Examination (OITE). Performance was similar for text-only and image-based questions when images were described by experts, but declined with AI-generated descriptions.
Area of Science:
- Medical Education Technology
- Artificial Intelligence in Healthcare
- Orthopaedic Surgery Training
Background:
- Artificial intelligence (AI) demonstrates potential in medical education by answering complex questions.
- Evaluating AI performance on specialized medical examinations is crucial for understanding its capabilities.
- The Orthopaedic In-Training Examination (OITE) serves as a benchmark for orthopaedic resident knowledge.
Purpose of the Study:
- To assess the performance of ChatGPT (GPT-4) on the complete 2019 Orthopaedic In-Training Examination (OITE).
- To compare ChatGPT's performance on text-only versus image-component questions.
- To investigate the impact of AI-generated versus human-expert image descriptions on ChatGPT's accuracy.
Main Methods:
- The 2019 OITE questions were inputted into ChatGPT version 4.0 (GPT-4).
- Image-based questions were supplemented with descriptions generated by Microsoft Azure AI Vision Studio or orthopaedic specialists.
- ChatGPT's responses were analyzed for accuracy across different question formats and image description methods.
Main Results:
- ChatGPT achieved an average correct answer rate of 49% for text-only and 48% for image-component questions.
- Performance decreased by 6% when using AI-generated image descriptions compared to specialist descriptions.
- ChatGPT's overall score (49%) was lower than all resident classes, including PGY-1 residents, on the 2019 OITE.
Conclusions:
- ChatGPT performed below the level of all resident classes on the 2019 OITE.
- AI accuracy on image-based questions is dependent on the quality of image descriptions, with specialist descriptions yielding better results.
- Understanding AI performance limitations is vital for its responsible integration into medical education and healthcare.

