Related Experiment Video
Updated: Aug 5, 2026

Orthopedic Robot-Assisted Femoral Neck System in the Treatment of Femoral Neck Fracture
Published on: March 3, 2023
Diagnostic Performance of Artificial Intelligence (AI) Chatbot Compared to Orthopedic Trauma Surgeons in Evaluating
Maria Oulianski1, Rami Mosheiff1, Dana Avraham1
1Hadassah Medical Center of the Hebrew University, Jerusalem, Israel.
Purpose Of The Study:
The indication for surgery for isolated lateral malleolar fracture (AO/OTA 44B1) is debatable and in many cases, relies upon radiographic assessment of fracture stability. Artificial intelligence chatbots with visual analysis capabilities offer a potential attribute for radiographic assessment. This study compared the diagnostic performance of a commercially available AI chatbot with that of three fellowship-trained orthopedic trauma surgeons in evaluating equivocal isolated lateral malleolar fractures.
Material And Methods:
A retrospective study. 50 patients with isolated lateral malleolar injury at the level of the syndesmosis (AO/OTA 44B1) were evaluated by three blinded fellowship-trained orthopedic trauma surgeons. Each rater measured standardized radiographic ankle parameters, medial clear space, tibiofibular clear space, and tibiofibular overlap, on anteroposterior and mortise views and determined a surgical versus nonoperative treatment recommendation. Subsequently, the same sets of radiographs were independently evaluated by an AI chatbot (Claude, Anthropic). The observers and AI decisions were compared to the actual outcome of the patients (operative vs. nonoperatives).
Results:
All raters recommended surgery at lower rates (34.0-46.0%) than the actual operative rate (56.0%). The difference in outcomes between the actual treatment and the observers varied and ranged between 67.3-86% with the AI within the same ranges. The AI's radiographic measurements differed systematically from all surgeons across five of six parameters. Inter-rater agreement between the AI and surgeons was slight, while inter-surgeon agreement was moderate (κ = 0.457-0.589). ROC analysis showed comparable AUC values (0.63-0.67) for all raters.
Discussion:
The AI chatbot demonstrated diagnostic accuracy comparable to orthopedic trauma surgeons in directing treatment for isolated lateral malleolar fractures, despite using a systematically different measurement strategy. All raters exhibited conservative bias as comparted with the actual outcome with modest discriminatory ability, reflecting the inherent difficulty of this clinical issue. These findings support a potential complementary role for AI in ankle fracture triage, while final clinical management decisions should remain in the hands of the orthopedic surgeon.