Related Experiment Videos
Human Versus Machine: Concordance and Conflict in Spinal Metastasis Management
Elie Najjar1, Rawan Masarwa1, Ahmed A Hassan1,2
1Centre for Spinal Studies and Surgery, Queens Medical Centre, Nottingham University Hospitals NHS Trust, Nottingham, UK.
Study Design:
Retrospective concordance analysis.
Objective:
To evaluate the concordance between ChatGPT-4 and a spinal multidisciplinary team (MDT) in management decisions for metastatic spinal disease and to explore the potential role of large language models (LLMs) as decision-support tools in oncologic spine care.
Summary Of Background Data:
Management of metastatic spinal disease requires multidisciplinary input balancing neurological preservation, mechanical stability, and systemic prognosis. Frameworks such as the Spine Instability Neoplastic Score (SINS) and the NOMS model guide decisions, but variability persists. Artificial intelligence tools like ChatGPT may standardize reasoning, yet their concordance with expert MDTs remains unexplored in this context.
Methods:
A retrospective analysis was performed on 100 anonymized adult cases referred to a spinal MDT for spinal metastases with suspected instability or metastatic spinal cord compression (MSCC). Ten additional cases were used to calibrate the model input. Each case summary, including demographics, tumor histology, SINS, Karnofsky Performance Scale, and Tokuhashi estimates, was presented identically to ChatGPT-4 (OpenAI, March 2025 snapshot) and the MDT. Management recommendations: surgical versus palliative (noncurative therapy including radiotherapy, systemic treatment, or best supportive care) were compared using Cohen κ, and discordant cases underwent qualitative analysis.
Results:
ChatGPT and the MDT were concordant in 86% of cases (κ=0.66, P<0.001), indicating substantial agreement. Concordant recommendations included surgery in 26% and palliative care in 60%. Discordance occurred in 14% of cases, typically involving younger patients or rare tumor histologies. ChatGPT emphasized mechanical instability and high SINS scores, whereas the MDT weighted systemic prognosis more heavily.
Conclusions:
ChatGPT-4 achieved substantial concordance with expert MDT decision-making in metastatic spinal disease. While it aligned well in structurally clear-cut cases, discordance in prognostically complex scenarios underscores the need for prognosis-aware, dynamically validated AI models. These findings support AI as an adjunct-not a replacement-for multidisciplinary clinical workflows.