Related Experiment Video
Updated: Sep 16, 2025

Author Spotlight: Fu's Subcutaneous Needling for Knee Osteoarthritis Pain
Published on: March 24, 2023
ChatGPT-4.0 and DeepSeek-R1 does not yet provide clinically supported answers for knee osteoarthritis
Haodong Wu1, Shuxin Yao2, Huanli Bao2
1Department of Knee Joint Surgery, Honghui Hospital, Xi'an Jiaotong University, No. 555 E.Youyi Rd, Xi'an 710061, China; Medical College of Yan'an University, Yan'an University, Yan'an 716000, China.
Background:
Large Language Models (LLMs) such as ChatGPT-4.0 and DeepSeek-R1 provide advanced natural language capabilities, but they also raise concerns regarding accuracy in medical applications. There is a lack of systematic evaluation of their performance against orthopedic guidelines, particularly for knee osteoarthritis (KOA). This study assessed the accuracy and consistency of these LLMs in relation to the most recent Chinese clinical practice guidelines for KOA.
Methods:
Queries regarding 17 guideline-recommended KOA therapeutic strategies were posed to ChatGPT-4.0 and DeepSeek-R1. Two independent reviewers evaluated response concordance (Concordance, Discordance, or No Concordance) with guidelines. Inter-rater reliability was assessed using Cohen's kappa coefficient. A chi-square test was employed to compare the response patterns between the two models.
Results:
ChatGPT-4.0 showed 59 % concordance; DeepSeek-R1 achieved 71 %. Both models gave inconsistent recommendations for ozone therapy and arthroscopy. ChatGPT-4.0 had five inconsistent responses; DeepSeek-R1 had three. Inter-rater agreement was high (κ = 0.90 and 0.86). No significant difference was found in concordance rates (P = 0.7; P = 1). Only DeepSeek-R1 provided references (38 in total), but just 8 were fully verifiable.
Conclusion:
Neither ChatGPT-4.0 nor DeepSeek-R1 consistently produced responses aligned with evidence-based clinical guidelines. These findings highlight the need for cautious interpretation of medical advice generated by current AI platforms, both by clinicians and patients.

