Related Experiment Video
Updated: Jan 17, 2026

Author Spotlight: Implementing the Enhanced Recovery After Surgery Concept in Rehabilitation Following Anterior Cruciate Ligament Reconstruction
Published on: March 1, 2024
ChatGPT-Generated Responses Across Orthopaedic Sports Medicine Surgery Vary in Accuracy, Quality, and Readability: A
Jacob D Kodra1, Arthur Saroyan2, Fabrizio Darby3
1Medical College of Wisconsin, Milwaukee, Wisconsin, U.S.A.
Purpose:
To evaluate the current literature regarding the accuracy and efficacy of ChatGPT in delivering patient education on common orthopaedic sports medicine operations.
Methods:
A systematic review was performed in accordance with Preferred Reporting Items for Systematic Reviews and Meta-analyses guidelines. After PROSPERO registration, a keyword search was conducted in the PubMed, Cochrane Central Register of Controlled Trials, and Scopus databases in September 2024. Articles were included if they evaluated ChatGPT's performance against established sources, examined ChatGPT's ability to provide counseling related to orthopaedic sports medicine operations, and assessed ChatGPT's quality of responses. Primary outcomes assessed were quality of written content (e.g., DISCERN score), readability (e.g., Flesch-Kincaid Grade Level and Flesch-Kincaid Reading Ease Score), and reliability (Journal of the American Medical Association Benchmark Criteria).
Results:
Seventeen articles satisfied the inclusion and exclusion criteria and formed the basis of this review. Four studies compared the effectiveness of ChatGPT and Google, and another study compared ChatGPT-3.5 with ChatGPT-4. ChatGPT provided moderate- to high-quality responses (mean DISCERN score, 41.0-62.1), with strong inter-rater reliability (0.72-0.91). Readability analyses showed that responses were written at a high school to college reading level (mean Flesch-Kincaid Grade Level, 10.3-16.0) and were generally difficult to read (mean Flesch-Kincaid Reading Ease Score, 28.1-48.0). ChatGPT frequently lacked source citations, resulting in a poor reliability score across all studies (mean Journal of the American Medical Association score, 0). Compared with Google, ChatGPT-4 generally provided higher-quality responses. ChatGPT also displayed limited source transparency unless specifically prompted for sources. ChatGPT-4 outperformed ChatGPT-3.5 in response quality (DISCERN score, 3.86 [95% confidence interval, 3.79-3.93] vs 3.46 [95% confidence interval, 3.40-3.54]; P = .01) and readability.
Conclusions:
ChatGPT provides generally satisfactory responses to patient questions regarding orthopaedic sports medicine operations. However, its utility remains limited by challenges with source attribution, high reading complexity, and variability in accuracy.
Level Of Evidence:
Level V, systematic review of Level V studies.

