Related Experiment Video
Updated: Mar 31, 2026

07:32
Standardized Histomorphometric Evaluation of Osteoarthritis in a Surgical Mouse Model
Published on: May 6, 2020
13.5K
Evaluating Large Language Model Adherence to AAOS Knee Osteoarthritis Guidelines: A Comparative Study of ChatGPT and
Cemil Yıldız1, Mehmet Yiğit Gökmen2, Çetin Utlu3
1Department of Orthopaedics and Traumatology, Gulhane Faculty of Medicine, University of Health Sciences, Ankara, Türkiye.
Indian Journal of Orthopaedics
|March 30, 2026
Summary
Large language models like ChatGPT and NotebookLM show strong adherence to orthopedic guidelines for knee osteoarthritis. AI reasoning aligns well with clinical practice, suggesting potential as educational tools.
Area of Science:
- Orthopedic Surgery
- Artificial Intelligence
- Medical Guidelines
Background:
- Clinical practice guidelines are essential for evidence-based orthopedic care.
- Large language models (LLMs) offer potential for synthesizing and disseminating medical information.
- Evaluating LLM alignment with established orthopedic guidelines is crucial for safe and effective implementation.
Purpose of the Study:
- To assess the concordance of LLM-generated reasoning with the American Academy of Orthopaedic Surgeons (AAOS) clinical practice guidelines for knee osteoarthritis (OA).
- To compare the adherence of ChatGPT and NotebookLM to specific AAOS guidelines for knee OA management.
Main Methods:
- A mixed-methods approach combining quantitative scoring and qualitative analysis.
- Thirty-three decision points from AAOS guidelines were used to create PICO prompts for LLMs.
- Orthopedic surgeons rated LLM responses on accuracy, evidence reasoning, and knowledge integration, with concordance classified as full, partial, or discordant.
Main Results:
- ChatGPT achieved a mean score of 3.67 and NotebookLM 3.55, with no significant difference (p=0.18).
- Full concordance was observed in 84.8% of ChatGPT responses and 75.8% of NotebookLM responses.
- Both models performed consistently in high-evidence areas, with variability in limited-evidence domains.
Conclusions:
- LLMs demonstrate substantial alignment with evidence-based orthopedic reasoning for knee OA.
- ChatGPT exhibited slightly higher fidelity to recommendation strength, while NotebookLM offered broader interpretation.
- Structured prompting can enhance LLM consistency, supporting their role in evidence translation and orthopedic education.

