Related Experiment Video
Updated: Jun 2, 2025

A Retrospective Study on Endoscopic Surgery for the Treatment of Paravertebral Abscess in Spinal Tuberculosis Patients
Published on: October 25, 2024
Comparative diagnostic accuracy of ChatGPT-4 and machine learning in differentiating spinal tuberculosis and spinal
Xiaojiang Hu1, Dongcheng Xu2, Hongqi Zhang3
1Department of Spine Surgery and Orthopaedics, Xiangya Hospital, Central South University, Changsha 410008, China; Department of Orthopedics, The Second Xiangya Hospital of Central South University, Changsha, 410011, Hunan, China.
Background:
In clinical practice, distinguishing between spinal tuberculosis (STB) and spinal tumors (ST) poses a significant diagnostic challenge. The application of AI-driven large language models (LLMs) shows great potential for improving the accuracy of this differential diagnosis.
Purpose:
To evaluate the performance of various machine learning models and ChatGPT-4 in distinguishing between STB and ST.
Study Design:
A retrospective cohort study.
Patient Sample:
A total of 143 STB cases and 153 ST cases admitted to Xiangya Hospital Central South University, from January 2016 to June 2023 were collected.
Outcome Measures:
This study incorporates basic patient information, standard laboratory results, serum tumor markers, and comprehensive imaging records, including Magnetic Resonance Imaging (MRI) and Computed Tomography (CT), for individuals diagnosed with STB and ST. Machine learning techniques and ChatGPT-4 were utilized to distinguish between STB and ST separately.
Method:
Six distinct machine learning models, along with ChatGPT-4, were employed to evaluate their differential diagnostic effectiveness.
Result:
Among the 6 machine learning models, the Gradient Boosting Machine (GBM) algorithm model demonstrated the highest differential diagnostic efficiency. In the training cohort, the GBM model achieved a sensitivity of 98.84% and a specificity of 100.00% in distinguishing STB from ST. In the testing cohort, its sensitivity was 98.25%, and specificity was 91.80%. ChatGPT-4 exhibited a sensitivity of 70.37% and a specificity of 90.65% for differential diagnosis. In single-question cases, ChatGPT-4's sensitivity and specificity were 71.67% and 92.55%, respectively, while in re-questioning cases, they were 44.44% and 76.92%.
Conclusion:
The GBM model demonstrates significant value in the differential diagnosis of STB and ST, whereas the diagnostic performance of ChatGPT-4 remains suboptimal.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Related Concept Videos
Pulmonary Tuberculosis IV
Several diagnostic approaches are used to detect TB. The conventional method is the Tuberculin Skin Test (TST), also known as the Mantoux test. However, this method has...
Pulmonary Tuberculosis III
The first classification is based on the development of the disease, and it includes the following categories: