Related Experiment Video
Updated: Aug 5, 2026

07:15
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
6.9K
Evaluating the Accuracy of Privacy-Preserving Large Language Models in Calculating the Spinal Instability Neoplastic
Li Yi Tammy Chan1, Ding Zhou Matthew Chan1, Yi Liang Tan1
1Department of Diagnostic Imaging, National University Hospital, Singapore 119074, Singapore.
Cancers
|July 12, 2025
Summary
Claude 3.5 accurately computed the Spine Instability Neoplastic Score (SINS) from radiology reports, showing high agreement with clinicians. This suggests large language models (LLMs) can aid in assessing spinal metastases.
Area of Science:
- Medical Imaging and Artificial Intelligence
- Oncology and Radiology
- Natural Language Processing in Healthcare
Background:
- Large language models (LLMs) are increasingly utilized in healthcare, with potential applications in diagnostic radiology.
- The Spine Instability Neoplastic Score (SINS) is crucial for evaluating spinal metastases, but LLM accuracy in its computation from radiological reports is not well-established.
Purpose of the Study:
- To assess the accuracy of two privacy-preserving LLMs, Claude 3.5 and Llama 3.1, in calculating the SINS using radiology reports and electronic medical records.
- To compare the performance of these LLMs against human clinician readers in SINS computation.
Main Methods:
- A retrospective analysis of 124 radiology reports from patients with spinal metastases.
- Independent SINS calculation by three expert readers (reference standard), two orthopaedic surgery residents, and two LLMs (Claude 3.5, Llama 3.1).
- Inter-rater agreement measured using intraclass correlation coefficient (ICC) for total SINS and Gwet's Kappa for individual components.
Main Results:
- Both LLMs and clinicians achieved near-perfect agreement with the reference standard for total SINS.
- Claude 3.5 (ICC = 0.984) demonstrated superior performance compared to Llama 3.1 (ICC = 0.829).
- Claude 3.5's performance was comparable to clinician readers (ICCs ranging from 0.926 to 0.986) across all SINS components.
Conclusions:
- Claude 3.5 exhibits high accuracy in calculating the SINS, indicating its potential as a valuable tool in clinical workflows.
- LLMs like Claude 3.5 may help reduce clinician workload and maintain diagnostic reliability in assessing spinal metastases.
- Further validation and optimization are necessary before widespread clinical integration of LLMs for SINS calculation due to observed performance variations.

