Related Experiment Video
Updated: Apr 15, 2026

The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve
Published on: September 7, 2022
ChatGPT provides accurate and safe responses to patient questions on hip arthroscopy, while completeness remains
Nikolai Ramadanov1,2, Plamen Penchev3, Maximilian Voss1,2
1Center of Orthopaedics and Traumatology, Brandenburg Medical School, University Hospital Brandenburg an der Havel, Brandenburg an der Havel, Germany.
Purpose:
Large language models, such as ChatGPT, are increasingly used by patients seeking information on hip arthroscopy (HAS) and femoroacetabular impingement (FAI). Despite their linguistic fluency, the accuracy, completeness and safety of procedure-specific patient information remain unclear. Although orthopaedic studies report variable performance across subspecialties, no systematic evaluation has specifically addressed HAS.
Methods:
PubMed, Embase, Scopus, CINAHL and Epistemonikos were searched to 10 February 2026 for studies evaluating ChatGPT responses to patient-oriented HAS or FAI questions. Randomized and non-randomized studies, observational cohorts and case series were eligible. Data on question sources, model versions, rating systems and performance domains (accuracy, relevance, completeness, safety, readability and clarity) were extracted. Heterogeneous rating scales were dichotomized into high- versus low-quality responses. Risk of bias was assessed using QUADAS-2 and ROBINS-I. Random-effects single-arm meta-analyses (REML) were conducted for each domain.
Results:
Eight studies met eligibility criteria. Accuracy was high (pooled 88.6%). Relevance, safety, readability and clarity reached pooled values of 100% with low heterogeneity. Completeness was lower (83.8%) with moderate heterogeneity, mainly driven by early GPT-3.5 studies. Funnel plots showed no clear small-study effects, although interpretation was limited by the small number of studies. Risk of bias was predominantly high or moderate, largely due to non-systematic question selection and heterogeneous rating tools. Later models (GPT-4/4o and beyond) demonstrated higher performance compared with GPT-3.5.
Conclusion:
ChatGPT provides accurate, relevant, safe and clear responses to patient questions about HAS, while completeness shows moderate variability. Although LLMs appear promising as adjuncts to patient education, methodological limitations in the current evidence base underscore the need for expert clinical counselling and more rigorous, standardized evaluation frameworks.
Level Of Evidence:
Level III systematic review and meta-analysis of non-randomized studies.

