Related Experiment Video
Updated: Jul 2, 2026

A Spine Robotic-Assisted Navigation System for Pedicle Screw Placement
Published on: May 11, 2020
Artificial intelligence fails to outperform orthopaedic surgeons: A systematic review
Jemima Russell1, Jamie Rosen1,2, Martinique Vella-Baldacchino1
1Department of Surgery and Cancer MSk Lab-Imperial College London London UK.
Purpose:
Artificial intelligence (AI) in orthopaedic surgery is increasingly applied to analyse clinical data, triage patients and interpret imaging with high accuracy. Orthopaedics surgery faces unique challenges, including high patient volumes, complex cases and prolonged waiting lists, highlighting the need for efficiency and decision support. To justify implementation, AI must demonstrate performance comparable to surgeons. This systematic review evaluates AI's performance relative to surgeons to determine its value as a complementary tool in orthopaedic practice.
Methods:
This systematic review was conducted using OVID Medline. Relevant studies published up to 13 August 2025 were identified. Included studies were categorised into decision making, management plans, clinical knowledge, quality control, and answering patients' frequently asked questions (FAQs).
Results:
Of 419 identified studies, 16 were eligible. ChatGPT showed high sensitivity in identifying patients achieving clinically meaningful improvements (97% vs. 90% for surgeons) but lower specificity (33% vs. 63%) and accuracy (65% vs. 76%). AI demonstrates comparable or superior performance to surgeons in emergency scenarios and answering patient FAQs, scoring higher across empathy, accuracy, completeness and overall quality (4.4 vs. 3.5-3.7). Residents outperformed AI in examinations (74.2% vs. 47.2%). AI showed limited accuracy in knee osteoarthritis radiographic staging (35% vs. >80%).
Conclusions:
AI demonstrates the potential to support clinical efficiency and patient communication in orthopaedics. However, concerns about bias, quality risks, overconfidence and reliance on outdated information prevent it from replacing human expertise. Clinician-led design and validation are required to ensure safe and effective integration into clinical practice.
Level Of Evidence:
Level IV.

