Related Experiment Video
Updated: Jan 11, 2026

Author Spotlight: Implementing the Enhanced Recovery After Surgery Concept in Rehabilitation Following Anterior Cruciate Ligament Reconstruction
Published on: March 1, 2024
How accurately do large language models answer patient questions on anterior cruciate ligament tears? A comparative
Salim Youssef1, Anton Brehmer1, Peter Melcher2
1University of Leipzig, Department of Orthopedics, Trauma and Plastic Surgery, Leipzig, Germany.
Background:
Large language models (LLMs) are increasingly used in the medical sector, raising questions about their reliability for patient education. With more LLMs becoming publicly available, it remains unclear whether meaningful performance differences exist between them. This is particularly relevant for anterior cruciate ligament (ACL) injuries, which mainly affect young, active individuals, those most likely to seek health advice from AI. This study aimed to evaluate and directly compare the accuracy of five leading LLMs in answering common patient questions about ACL tears.
Methods:
Fourteen commonly asked patient questions were identified in a systematic online search. Each question was submitted to five LLMs: ChatGPT-4, Gemini 2.0, Llama 3.1, DeepSeek-V3, and Grok3. Responses were assessed for accuracy by orthopedic consultants using a five-point Likert scale. Word count was recorded as a proxy for readability. Statistical analysis included ANOVA by Tukey's HSD post hoc test.
Results:
All models achieved mean accuracy scores ≥3 (mostly accurate). DeepSeek (3.61) and Grok (3.59) demonstrated significantly higher mean accuracies than Llama (3.25; P < 0.05). ChatGPT and Gemini achieved mean scores of 3.48 and 3.52, respectively. Models generating longer responses, such as Grok and DeepSeek, tended to offer greater accuracy, whereas Llama produced the shortest and least accurate answers.
Conclusions:
All tested LLMs show promise for patient education regarding ACL injuries, but notable performance differences exist. Model choice is therefore critical. While all responses were evaluated by clinical experts, the lack of guideline-based validation highlights the need for further studies assessing both accuracy and patient comprehension.
More Related Videos
06:28Anterior Cruciate Ligament Transection and Synovial Fluid Lavage in a Rodent Model to Study Joint Inflammation and Posttraumatic Osteoarthritis
Published on: September 2, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024