Related Experiment Video
Updated: Jul 19, 2026

06:09
Measuring 3D In-vivo Shoulder Kinematics using Biplanar Videoradiography
Published on: March 12, 2021
3.0K
Assessment and comparison of artificial intelligence-generated information regarding shoulder arthroplasty from
Suhasini Gupta1, Brett D Haislup2, Anisha Tyagi3
1T.H. Chan School of Medicine, University of Massachusetts, Worcester, MA, USA.
Journal of Shoulder and Elbow Surgery
|February 19, 2025
Summary
Microsoft CoPilot provided more reliable and easier-to-read information on anatomic total shoulder arthroplasty (aTSA) and reverse total shoulder arthroplasty (rTSA) than ChatGPT. However, both AI tools should supplement, not replace, primary patient resources.
Area of Science:
- Artificial Intelligence in Medicine
- Orthopedic Surgery Information Dissemination
- Patient Education Technology
Background:
- Evaluating the quality, accuracy, and readability of AI-generated information on shoulder arthroplasty is crucial for patient understanding.
- Anatomic total shoulder arthroplasty (aTSA) and reverse total shoulder arthroplasty (rTSA) are common orthopedic procedures with distinct indications.
- Patient-facing information from AI tools like ChatGPT and CoPilot needs rigorous assessment for clinical use.
Purpose of the Study:
- To compare the quality, accuracy, and readability of AI-generated responses for aTSA and rTSA patient queries.
- To assess information provided by OpenAI's ChatGPT and Microsoft's CoPilot using established medical and readability benchmarks.
- To analyze the reliability and source diversity of AI-generated medical information.
Main Methods:
- Thirty common patient questions on aTSA and rTSA were posed to ChatGPT 3.5 and CoPilot.
- Responses were evaluated using DISCERN, JAMA benchmark criteria, Flesch-Kincaid Reading Ease Score (FRES), and Flesch-Kincaid Grade Level (FKGL).
- CoPilot's citation sources were analyzed for academic and medical relevance.
Main Results:
- Both AI interfaces provided 'good' quality information (DISCERN >50), with CoPilot scoring higher overall.
- CoPilot demonstrated superior reliability (higher JAMA score) and readability (higher FRES) compared to ChatGPT.
- CoPilot provided citations, with a notable percentage of academic sources, while ChatGPT's responses were more complex and academic in reading level.
Conclusions:
- CoPilot offers more reliable, accessible, and better-referenced information for shoulder arthroplasty queries than ChatGPT.
- Despite improved readability, CoPilot's 12th-grade reading level may still challenge patient comprehension.
- AI-generated content, including CoPilot's, should be used as supplementary resources, not primary sources, for patient education on shoulder arthroplasty.
Keywords:
Total shoulder arthroplastyartificial intelligencehealth literacypatient educationreverse shoulder arthroplastytechnology in shoulder arthroplasty
