Related Experiment Video
Updated: Jun 20, 2026

Treatment of Osteochondral Defects in the Rabbit's Knee Joint by Implantation of Allogeneic Mesenchymal Stem Cells in Fibrin Clots
Published on: May 21, 2013
Performance of Artificial Intelligence in Addressing Questions Regarding Management of Osteochondritis Dissecans
John D Milner1, Matthew S Quinn1, Phillip Schmitt1
1Department of Orthopaedic Surgery, Brown University, Warren Alpert Medical School, Providence, Rhode Island.
Background:
Large language model (LLM)-based artificial intelligence (AI) chatbots, such as ChatGPT and Gemini, have become widespread sources of information. Few studies have evaluated LLM responses to questions about orthopaedic conditions, especially osteochondritis dissecans (OCD).
Hypothesis:
ChatGPT and Gemini will generate accurate responses that align with American Academy of Orthopaedic Surgeons (AAOS) clinical practice guidelines.
Study Design:
Cohort study.
Level Of Evidence:
Level 2.
Methods:
LLM prompts were created based on AAOS clinical guidelines on OCD diagnosis and treatment, and responses from ChatGPT and Gemini were collected. Seven fellowship-trained orthopaedic surgeons evaluated LLM responses on a 5-point Likert scale, based on 6 categories: relevance, accuracy, clarity, completeness, evidence-based, and consistency.
Results:
ChatGPT and Gemini exhibited strong performance across all criteria. ChatGPT mean scores were highest for clarity (4.771 ± 0.141 [mean ± SD]). Gemini scored highest for relevance and accuracy (4.286 ± 0.296, 4.286 ± 0.273). For both LLMs, the lowest scores were for evidence-based responses (ChatGPT, 3.857 ± 0.352; Gemini, 3.743 ± 0.353). For all other categories, ChatGPT mean scores were higher than Gemini scores. The consistency of responses between the 2 LLMs was rated at an overall mean of 3.486 ± 0.371. Inter-rater reliability ranged from 0.4 to 0.67 (mean, 0.59) and was highest (0.67) in the accuracy category and lowest (0.4) in the consistency category.
Conclusion:
LLM performance emphasizes the potential for gathering clinically relevant and accurate answers to questions regarding the diagnosis and treatment of OCD and suggests that ChatGPT may be a better model for this purpose than the Gemini model. Further evaluation of LLM information regarding other orthopaedic procedures and conditions may be necessary before LLMs can be recommended as an accurate source of orthopaedic information.
Clinical Relevance:
Little is known about the ability of AI to provide answers regarding OCD.

