Related Experiment Video
Updated: Mar 31, 2026

In Vivo Quantification of Hip Arthrokinematics during Dynamic Weight-bearing Activities using Dual Fluoroscopy
Published on: July 2, 2021
Evaluating ChatGPT-5's Performance in Answering Common Patient Questions About Femoroacetabular Impingement and Hip
Maximilian Voss1,2, Hannah Jaeger1,2, Mikhail Salzmann1,2
1Center of Orthopaedics and Traumatology, Brandenburg Medical School, University Hospital Brandenburg/Havel, Brandenburg an der Havel, Germany.
Background:
Hip arthroscopy (HAS) is widely used to treat femoroacetabular impingement syndrome (FAIS), and many patients rely on online resources for medical information. Large language models (LLMs) such as ChatGPT have shown potential as supplementary educational tools in orthopedics; however, existing evaluations are limited to earlier model generations with variable accuracy and completeness. This study aimed to evaluate the accuracy, clarity, relevance, and completeness of ChatGPT-5 responses to common patient questions regarding FAIS and HAS.
Methods:
ChatGPT-5 was used to generate 25 frequently asked patient questions and corresponding answers related to hip preservation. Two fellowship-trained hip preservation surgeons independently evaluated each response using a five-point Likert scale across four predefined domains: relevance, accuracy, clarity, and completeness. Descriptive statistics were calculated as mean ± standard deviation for each domain. Inter-rater reliability was assessed using a two-way random-effects intraclass correlation coefficient with absolute agreement (ICC [2, 1]) and complemented by exact agreement percentages.
Results:
All responses received excellent scores, with mean values ranging from 4.84 ± 0.27 (completeness) to 5.00 ± 0.00 (relevance). Accuracy (4.97 ± 0.08) and clarity (4.91 ± 0.17) were near-perfect. ICC values demonstrated moderate to excellent agreement (0.70-0.81), complemented by high exact agreement rates (84-100%). No answer contained factually incorrect, misleading, or unsafe information. Minor reductions in completeness were attributable to occasional brevity rather than substantive omissions.
Conclusion:
ChatGPT-5 generated highly accurate, clear, and clinically appropriate patient-oriented explanations regarding FAIS and HAS, showing clear improvement compared with earlier ChatGPT versions. Although ChatGPT-5 represents a marked advancement in AI-based patient education, its use should be regarded as a complementary educational tool rather than a replacement for professional orthopedic counseling.
More Related Videos
09:51The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve
Published on: September 7, 2022
06:16A Probing Device for Quantitatively Measuring the Mechanical Properties of Soft Tissues during Arthroscopy
Published on: May 1, 2020