Related Experiment Video
Updated: Jun 6, 2025

05:57
A Teleoperated Robotic System-Assisted Percutaneous Transiliac-Transsacral Screw Fixation Technique
Published on: January 6, 2023
2.2K
Exploring the Performance of ChatGPT in an Orthopaedic Setting and Its Potential Use as an Educational Tool
Arthur Drouaud1, Carolina Stocchi2, Justin Tang2
1George Washington University School of Medicine, Washington, District of Columbia.
JB & JS Open Access
|November 27, 2024
Summary
ChatGPT-4 vision (GPT-4V) shows fair agreement in medical reasoning but falls below surgeon standards for orthopaedic trauma education, particularly in image interpretation. Its performance in management and treatment is stronger than in imaging analysis.
Area of Science:
- Artificial Intelligence in Medical Education
- Large Language Models in Healthcare
- Orthopaedic Trauma Training
Background:
- The integration of artificial intelligence (AI) tools like ChatGPT-4 vision (GPT-4V) into medical education is rapidly evolving.
- Evaluating AI's capabilities in interpreting medical data and formulating patient management strategies is crucial for its effective implementation.
- Orthopaedic trauma education presents complex diagnostic and management challenges where AI could potentially assist.
Purpose of the Study:
- To assess the performance of GPT-4V in interpreting orthopaedic trauma medical images and patient information.
- To evaluate GPT-4V's capabilities in diagnosis formulation and guiding patient management and treatment decisions.
- To determine the potential of GPT-4V as an educational tool for medical students in real-life orthopaedic trauma cases.
Main Methods:
- Ten popular orthopaedic trauma cases from OrthoBullets were selected for analysis.
- GPT-4V interpreted provided medical imaging and patient data, generating diagnoses and answering case-specific questions.
- Four fellowship-trained orthopaedic trauma surgeons rated GPT-4V's responses on accuracy, rationale, relevance, and trustworthiness using a 5-point Likert scale.
Main Results:
- GPT-4V received an overall mean rating of 3.46/5.00 for imaging interpretation, with specific scores for accuracy (3.28), rationale (3.68), relevance (3.75), and trustworthiness (3.15).
- Performance improved for management questions (overall 3.76) and treatment questions (overall 4.04), indicating better AI capabilities in these areas.
- Surgeon ratings suggested fair agreement with GPT-4V's reasoning but highlighted lower performance in imaging interpretation compared to management and treatment.
Conclusions:
- This study is the first to evaluate GPT-4V for orthopaedic trauma medical education, covering imaging, management, and treatment.
- GPT-4V demonstrated fair agreement in reasoning but did not meet the standards of fellowship-trained surgeons as a standalone educational tool.
- The AI tool's performance was notably weaker in interpreting medical images compared to its effectiveness in providing management and treatment guidance.

