Related Experiment Video
Updated: Jan 17, 2026

Vascularized Composite Hand Allograft Procurement and Preparation for Distal and Proximal Forearm Allotransplantation: A Stepwise Approach
Published on: May 23, 2025
Do ChatGPT and Gemini's Recommendations Align With Established Guidelines for Hand and Upper Extremity Surgery?
Yibin B Zhang1, Fielding S Fischer1, Matthew V Abola2
1Harvard Medical School, Boston, MA, USA.
Large language models (LLMs) like ChatGPT and Gemini show moderate clinical accuracy for orthopedic guidelines. Both LLMs have limitations in transparency and consistency, requiring further evaluation before widespread clinical adoption.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Healthcare Informatics
Background:
- Large language models (LLMs) are increasingly used in healthcare for administrative tasks and patient communication.
- Concerns exist regarding the clinical accuracy and reliability of LLM recommendations in medical practice.
- This study investigates the adherence of LLM outputs to established clinical practice guidelines (CPGs).
Purpose of the Study:
- To evaluate the concordance of recommendations from ChatGPT and Gemini with American Academy of Orthopedic Surgeons (AAOS) clinical practice guidelines (CPGs).
- To compare the performance of ChatGPT and Gemini in terms of accuracy and citation transparency across different orthopedic conditions.
Main Methods:
- ChatGPT (version 4o) and Gemini (version 1.5 Flash) were prompted with text aligned to AAOS CPGs for carpal tunnel syndrome, distal radius fractures, and glenohumeral joint osteoarthritis.
- Blinded reviewers assessed the LLM outputs for concordance with CPGs.
- Concordance rates were compared between models, topics, and guideline strengths, with citation transparency also evaluated.
Main Results:
- An overall concordance rate of 62.1% was observed across 174 recommendations.
- No statistically significant difference in concordance was found between ChatGPT (66.7%) and Gemini (57.5%) (P = .131).
- Both models exhibited low citation transparency, with Gemini providing sources more frequently (39.1%) than ChatGPT (3.5%) (P < .0001).
Conclusions:
- Current LLMs demonstrate modest concordance with AAOS CPGs, indicating significant limitations for clinical integration.
- Variability in performance across topics and guideline strengths, coupled with poor citation transparency, necessitates further refinement.
- Thorough evaluation and improvement are crucial before adopting LLMs in hand surgery and other clinical specialties.
More Related Videos
09:14Surface Electromyographic Biofeedback as a Rehabilitation Tool for Patients with Global Brachial Plexus Injury Receiving Bionic Reconstruction
Published on: September 28, 2019
07:59Therapy Interventions for Upper Limb Amputees Undergoing Selective Nerve Transfers
Published on: October 29, 2021
Related Concept Videos
Pre-Procedural Guidelines for Assessing Blood Pressure
Peripheral Artery Disease V: Postoperative Nursing Management