Related Experiment Video
Updated: Jun 28, 2026

Autologous Microfractured and Purified Adipose Tissue for Arthroscopic Management of Osteochondral Lesions of the Talus
Published on: January 23, 2018
Concordance of ChatGPT, Gemini, Claude, and OpenEvidence with the 2024 AAOS guidelines on acute isolated meniscal
Wei-Kuo Hsu1, Hao-Chun Chuang2, Yang-Yi Wang3
1Department of Orthopaedic Surgery, National Cheng Kung University Hospital, College of Medicine, National Cheng Kung University, Tainan, Taiwan; Department of Orthopaedic Surgery, National Cheng Kung University Hospital Dou-Liou Branch, College of Medicine, National Cheng Kung University, Yunlin, Taiwan.
Purpose:
To evaluate the reliability and clinical applicability of the three most commonly used large language models (LLMs) (ChatGPT, Gemini, and Claude) and a domain-specific artificial intelligence (AI) platform (OpenEvidence) in providing recommendations for acute isolated meniscal pathology, to compare their accuracy, and to assess the consistency between American Academy of Orthopedic Surgeons (AAOS) Clinical Practice Guidelines (CPG) recommendations and AI-generated guidance.
Methods:
An exploratory cross-sectional benchmarking analysis evaluated concordance of three large language models (ChatGPT, Gemini, Claude) and one domain-specific AI (OpenEvidence) with 2024 AAOS clinical practice guidelines for acute isolated meniscal pathology. Nine guideline recommendations were converted into standardized questions and presented to each AI model on the same day. Three sports medicine orthopedic specialists independently assessed responses as concordant or discordant, with disagreements resolved by majority decision. Statistical analysis used SPSS 29, employing Cochran's Q test for concordance assessment and Fleiss' kappa for inter-rater reliability.
Results:
OpenEvidence achieved perfect concordance (9/9, 100%), followed by ChatGPT (8/9, 89%), Gemini and Claude (both 7/9, 78%). Overall concordance rate was 86% (31/36). Concordance was 100% for strong and consensus recommendations, 75% for moderate and limited recommendations. Cochran's Q test showed no significant difference among models (Q = 3.00, p = 0.392). Inter-rater reliability demonstrated almost perfect agreement (κ = 0.825, 95% CI: 0.637-1.014).
Conclusions:
Although ChatGPT, Gemini, and Claude demonstrated high concordance with the AAOS CPG for acute isolated meniscal pathology, their responses were not consistently guideline-concordant. OpenEvidence achieved the highest descriptive concordance rate (100%); however, statistical superiority could not be established due to the limited number of guideline items. This exploratory benchmarking analysis suggests that domain-specific AI models may represent a valuable tool for retrieving information on acute isolated meniscal injuries.
Clinical Relevance:
The difference between domain-specific AI model and general LLMs underscore the need to educate the general public and clinicians about the limitations of general-purpose chatbots, emphasizing that LLM outputs should be interpreted with caution in real-world practice, while tools like OpenEvidence exist for evidence-based information.
More Related Videos
07:06Destabilization of the Medial Meniscus and Cartilage Scratch Murine Model of Accelerated Osteoarthritis
Published on: July 6, 2022
06:28Anterior Cruciate Ligament Transection and Synovial Fluid Lavage in a Rodent Model to Study Joint Inflammation and Posttraumatic Osteoarthritis
Published on: September 2, 2025