Related Experiment Video
Updated: Jun 21, 2025

10:07
Measurement of Dynamic Scapular Kinematics Using an Acromion Marker Cluster to Minimize Skin Movement Artifact
Published on: February 10, 2015
19.2K
Responses From ChatGPT-4 Show Limited Correlation With Expert Consensus Statement on Anterior Shoulder Instability
Alexander Artamonov1, Ira Bachar-Avnieli1,2, Eyal Klang3,4
1Orthopedic Department, Barzilai Medical Center, Ashkelon, Israel.
Arthroscopy, Sports Medicine, and Rehabilitation
|July 15, 2024
Summary
Generative Pretrained Transformer-4 (GPT-4) answers showed limited correlation with expert consensus on anterior shoulder instability (ASI) diagnosis and treatment. This highlights the need to verify AI-generated medical information against established clinical guidelines.
Area of Science:
- Orthopedics
- Artificial Intelligence in Medicine
- Medical Informatics
Background:
- Anterior shoulder instability (ASI) is a common orthopedic condition requiring accurate diagnosis and management.
- Expert consensus statements provide evidence-based guidelines for clinical practice.
- The increasing use of artificial intelligence (AI) in healthcare necessitates evaluating its alignment with established medical knowledge.
Purpose of the Study:
- To compare the accuracy and similarity of answers generated by Generative Pretrained Transformer-4 (GPT-4) against an expert consensus statement on anterior shoulder instability (ASI).
- To assess the reliability of AI in providing information consistent with current clinical guidelines for ASI.
Main Methods:
- An expert consensus statement on ASI diagnosis, nonoperative management, and Bankart repair was reviewed.
- GPT-4 was queried using the same questions posed to the expert panel.
- Answers from GPT-4 and the consensus statement were compared and rated for similarity by experienced shoulder surgeons.
- Interobserver reliability was calculated using weighted kappa scores.
Main Results:
- Shoulder surgeons rated GPT-4's responses as highly similar to the consensus statement for 25.8% of questions, medium for 45.2%, and low for 29%.
- GPT-4 self-assessed its responses as high similarity for 48.3%, medium for 41.9%, and low for 9.7%.
- Surgeons and GPT-4 agreed on similarity classification for 58.1% of questions, with disagreement on 41.9%.
Conclusions:
- AI-generated responses demonstrate a limited correlation with expert consensus statements regarding ASI diagnosis and treatment.
- The findings underscore the importance of critical evaluation and verification of AI-generated medical information.
- Further research is needed to refine AI models for clinical applications in orthopedics.

