Related Experiment Video
Updated: Jan 28, 2026

Reverse Total Shoulder Arthroplasty
Published on: July 5, 2011
Widely Available Large Language Models Are Not a Reliable Source to Address Medical Treatment Recommendations of
Elian Niklas Oudintsov1, Soraya Bahlawane1, Agahan Hayta1
1Department of Shoulder and Elbow Surgery, Center for Musculoskeletal Surgery, Campus Virchow Klinikum, Charité - Universitätsmedizin Berlin, corporate member of Freie Universität Berlin and Humboldt-Universität zu Berlin, Berlin, Germany.
Purpose:
To assess the ability of ChatGPT 3.5 to aid in the treatment planning process of first-time anteroinferior shoulder dislocation.
Methods:
Forty fictional patient cases were created varying in 15 different characteristics, whose distribution was randomized. Six orthopaedic surgeons (3 residents and 3 specialists in shoulder surgery) were then asked to determine the best treatment option for these patient cases. Their answers were compared with the treatment recommendations proposed by ChatGPT in 2 different sessions on the basis of preselected literature. To counteract the wide dispersion of responses, tendencies towards nonoperative, open surgical, or arthroscopic treatment were subsequently defined. The results were then analyzed descriptively.
Results:
The mean age of the fictional patients was 44 years (13-80 years), with 57.5% of the patients female. The agreement between the ChatGPT responses in the 2 sessions was 70.0%. In contrast, the 3 assistant physicians agreed with each other in 35% of all cases and the 3 specialists agreed in 32.5% of all cases. There was an exact match of 12.5% between the ChatGPT responses and all human assessments. In 65.0% of all cases, the physicians showed similar tendencies in their choice of therapy resulting in a 55.0% match between ChatGPT and the surgeons.
Conclusions:
There was no clear consensus regarding the treatment for first-time anteroinferior dislocations of the shoulder, neither among physicians nor with ChatGPT 3.5. However, ChatGPT 3.5 and physicians showed similar tendencies regarding the treatment in over half of the cases. Because of the inconsistent responses of ChatGPT 3.5, it cannot yet be considered as reliable tool for therapy planning.
Clinical Relevance:
ChatGPT 3.5, widely available and free of charge, is increasingly used in clinical settings. However, it's crucial to highlight its limitations in treatment planning for pathologies, especially when there's no clear consensus even among experienced surgeons.
Related Concept Videos
Reliability and Validity
Muscles of the Shoulder
Anterior Thoracic Muscles
The anterior thoracic muscles include the serratus anterior, subclavius, and...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

