Guideline Adherence and Variability in AI-Assisted Surgical Decision-Making for Low Back Pain
1Department of Orthopaedics and Traumatology, Başakşehir Çam and Sakura City Hospital, Istanbul, Türkiye.
Background:
Surgical decision-making in low back pain requires balancing diagnostic uncertainty, asymmetric risk and guideline adherence, particularly when distinguishing non-specific conditions from pathologies warranting timely surgical referral. As large language models (LLMs) are increasingly considered as decision-support tools, their reliability in guideline-concordant surgical triage requires systematic evaluation.
Objective:
To assess and compare the performance of two LLMs (GPT and Gemini) in guideline-adherent surgical decision-making across a range of lumbar spine presentations.
Methods:
This comparative cross-sectional study was conducted between January and February 2026. Sixty standardised clinical scenarios representing non-specific low back pain, lumbar disc herniation with radiculopathy, lumbar spinal stenosis, and red-flag emergency conditions were developed. Guideline-based management decisions served as the reference standard. Both models were queried using an identical fixed prompt to determine diagnosis, management strategy, and urgency. Sensitivity, specificity, overall accuracy, and area under the receiver operating characteristic (ROC) curve (AUC) were calculated. Discriminative performance was compared using the DeLong test.
Results:
Both models demonstrated high adherence in clearly defined emergency scenarios. GPT showed higher sensitivity for identifying surgical indications than Gemini (95.8% versus 75%), while specificity was 100% for both. Overall accuracy was 98.3% for GPT and 90% for Gemini. ROC analysis confirmed superior discriminative performance for GPT (AUC 0.99 versus 0.88; p < 0.05). Differences were primarily observed in non-emergent surgical cases.
Conclusions:
LLMs demonstrated strong alignment with guideline-based management in unambiguous scenarios, but variability emerged in sensitivity for non-emergent surgical indications. These findings underscore the importance of structured validation when integrating AI-assisted tools into clinical decision-making pathways.
