Related Experiment Video
Updated: Aug 14, 2026

10:40
Uniportal Full Endoscopic Posterolateral Transforaminal Lumbar Interbody Fusion
Published on: June 6, 2025
Large language models for lumbar spondylolisthesis detection: a multi-center pilot comparative radiographic accuracy
Rehan R Khan1,2, Tony Tannoury1,2, Rohith Ryali1,2
1Department of Orthopedic Surgery, Boston Medical Center, Boston, USA.
Summary
ChatGPT-5 shows improved performance over ChatGPT-4o in detecting lumbar spondylolisthesis on radiographs. However, both large language models (LLMs) still lag behind expert spine surgeons in diagnostic accuracy for this spinal condition.
Area of Science:
- Radiology
- Artificial Intelligence
- Spine Surgery
Background:
- Lumbar spondylolisthesis diagnosis relies on radiographic interpretation.
- Large language models (LLMs) show promise in medical image analysis, prompting patient interest in their diagnostic capabilities.
- The diagnostic reliability of LLMs for spinal pathologies is not well-established.
Purpose of the Study:
- To evaluate the performance of ChatGPT-4o and ChatGPT-5 in detecting lumbar spondylolisthesis.
- To compare LLM diagnostic accuracy against expert spine surgeon consensus.
- To assess the utility of advanced LLMs in interpreting standing lateral lumbar radiographs.
Main Methods:
- 200 standing lateral lumbar radiographs from the VinDr-SpineXR dataset were analyzed.
- 100 radiographs were positive and 100 were negative for spondylolisthesis.
- Five spine surgeons established expert consensus, and LLMs (ChatGPT-4o, ChatGPT-5) provided independent assessments using a binary prompt.
Main Results:
- Expert consensus confirmed spondylolisthesis in 81% of initially positive cases and reclassified 2% of negative cases as positive.
- ChatGPT-5 demonstrated higher sensitivity (67.5%) and accuracy (61.5%) compared to ChatGPT-4o (49.4% sensitivity, 55.5% accuracy).
- ChatGPT-5 showed better inter-rater agreement (κ=0.238) than ChatGPT-4o (κ=0.091).
Conclusions:
- ChatGPT-5 exhibits superior performance in identifying lumbar spondylolisthesis compared to ChatGPT-4o.
- Despite advancements, both LLMs exhibit limitations when compared to the diagnostic capabilities of fellowship-trained spine surgeons.
- Further research is needed to enhance LLM accuracy for spinal condition diagnosis.