Related Experiment Video
Updated: Jan 6, 2026

Automated Midline Shift and Intracranial Pressure Estimation based on Brain CT Images
Published on: April 13, 2013
In-context learning enables large language models to achieve human-level performance in spinal instability neoplastic
Maximilian F Russe1, Marco Reisert2,3, Anna Fink1
1Department of Diagnostic and Interventional Radiology, Faculty of Medicine, Medical Center - University of Freiburg, University of Freiburg, 79106, Freiburg, Germany.
Large language models (LLMs) can achieve near-human performance in classifying vertebral metastasis stability using the Spinal Instability Neoplastic Score (SINS) after task-specific refinement. This refinement, particularly in-context learning, is crucial for AI accuracy in medical applications.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Analysis
- Spinal Oncology
Background:
- Vertebral metastasis stability assessment is critical for patient management.
- The Spinal Instability Neoplastic Score (SINS) is a standard tool for this assessment.
- Evaluating the performance of AI models in this domain is essential.
Purpose of the Study:
- To compare the performance of state-of-the-art large language models (LLMs) against human experts in SINS classification.
- To assess the impact of task-specific refinement, including in-context learning, on LLM performance for SINS scoring.
- To explore the potential of AI in automating vertebral metastasis stability assessment.
Main Methods:
- Retrospective analysis of 100 synthetic CT and MRI reports with varying SINS scores.
- Comparison of performance between four human experts (radiologists, neurosurgeons) and four LLMs (Mistral, Claude, GPT-4 turbo, GPT-4o).
- LLMs were evaluated in both generic and task-specific refined forms, focusing on SINS category and point assignment accuracy.
Main Results:
- Human experts achieved high accuracy in SINS classification (98.5%) and point calculation (92%).
- Generic LLMs showed limited performance (26-63% correct classification, 4-15% correct SINS points).
- In-context learning significantly improved LLM performance to near-human levels (96-98% classification, 86-95% scoring), with refined models showing 71-85% improvement in SINS points allocation.
Conclusions:
- Task-specific refinement, especially in-context learning, enables LLMs to achieve near-human expert performance in SINS classification.
- LLMs show promise for automating vertebral metastasis stability assessment.
- The study underscores the necessity of task-specific refinement for effective AI implementation in medical contexts.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
