Related Experiment Video
Updated: Sep 6, 2026

A Postoperative Evaluation Guideline for Computer-Assisted Reconstruction of the Mandible
Published on: January 28, 2020
A Comparison of the Agreement of Four AI Models on Surgical Recommendations of MIRCTs Using the ASES Neer Circle
Cameron Vauclin1, Brandon Flaig1, Andrew Haynes2
1Louisiana State University Health Sciences Center Shreveport, Department of Orthopaedic Surgery, Shreveport, LA, USA.
Background:
With the rising utilization of large language models (LLMs), such as OpenEvidence (OE), ChatGPT-4o (GPT), Google Gemini (GG), and DeepSeek (DS), their use in complex Orthopaedic cases remains unclear. Massive irreparable rotator cuff tears (MIRCTs) represent a challenging clinical scenario requiring nuanced, patient-specific decision-making. We evaluated the concordance of LLM-generated surgical recommendations with expert consensus derived from a Delphi study by the American Shoulder and Elbow Surgeons (ASES) Neer Circle and compared the various LLM models against each other.
Methods:
Sixty-one MIRCT Delphi consensus scenarios were entered into the most current free version of each LLM in a standardized prompt format in June 2025. Recommendations were categorized as fully concordant or discordant with Delphi consensus. Further LLM testing included the addition of diabetes, smoking, and their combination. Accuracy (%) for each LLM and test with differences assessed using Cochran's Q test and McNemar's tests. A generalized linear mixed model identified significant predictors of AI-LLM accuracy, while the relationship between Delphi consensus strength and LLM concordance was assessed using Spearman's rank correlation coefficient.
Results:
A total of 976 recommendations (61 scenarios × 4 platforms x 4 tests) were analyzed. OE demonstrated the greatest accuracy across all 4 tests (65.6%, 68.9%, 62.3%, and 65.6%, respectively) (p < 0.05), while DS consistently scoring the lowest. Accuracy was significantly positively predicted by age greater than 70 (OR 31.6, p < 0.001), dynamic instability (OR 18.8, p = 0.002), and pseudoparesis (OR 2.9, p = 0.025), and negatively predicted by intact or anatomically reparable subscapularis (OR 0.17, p = 0.001). There was a significant positive correlation between the strength of Delphi expert consensus and the number of LLM platforms concordant with that consensus (Spearman's ρ = 0.493, p < 0.001).
Discussion & Conclusion:
This study is the first to systematically compare multiple LLMs against recommendations of an MIRCT Delphi consensus study. OE and GPT demonstrated the highest concordance; however, they did not approach levels of expert decision-making in many scenarios. LLMs have potential as adjunctive decision-support tools, particularly in resource-limited settings or for generalists managing complex shoulder pathology.
Level Of Evidence:
Basic Science Study, Computer Modeling using AI.
