Related Experiment Videos
Benchmarking 54 large language model configurations for CAD-RADS scoring: open-weight models approach human-level
Veit Sandfort1,2, Davis M Vigneault3, Martin J Willemink4
1Department of Radiology, Stanford University School of Medicine, 300 Pasteur Drive, Stanford, 94305, California, USA. veit.sandfort@gmail.com.
The International Journal of Cardiovascular Imaging
|July 30, 2026
Summary
Large language models (LLMs) can now classify Coronary Artery Disease Reporting and Data System (CAD-RADS) categories in coronary CT angiography reports. An open-weight LLM achieved performance comparable to proprietary models, even on consumer hardware.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing for Healthcare
- Cardiovascular Radiology Reporting Standards
Background:
- The Coronary Artery Disease Reporting and Data System (CAD-RADS) standardizes reporting for coronary CT angiography (CCTA).
- However, not all CCTA reports include CAD-RADS classifications, necessitating automated methods for extraction.
- Large language models (LLMs) show promise in processing unstructured medical text.
Purpose of the Study:
- To evaluate the zero-shot performance of various LLMs in classifying CAD-RADS categories from unstructured CCTA reports.
- To compare the performance of different LLM configurations, including proprietary and open-weight models.
- To assess the feasibility of using LLMs for automated CAD-RADS classification in clinical settings.
Main Methods:
- Retrospective analysis of 500 anonymized CCTA reports from four hospitals.
- Benchmarking of 54 LLM configurations (50 distinct models) using zero-shot prompts.
- Performance measured by unweighted Cohen's kappa (κ) against expert cardiovascular radiologist consensus.
Main Results:
- Two LLMs, Claude 4.6 Opus and open-weight Gemma 4 31B, met pre-specified non-inferiority criteria against human inter-rater agreement (κ > 0.81).
- Gemma 4 31B demonstrated robust performance, matching top proprietary systems, and is compatible with 3-bit quantization on a 24 GB consumer GPU.
- LLM performance was affected by report complexity (longer thinking chains), with 'thinking mode' outperforming 'non-thinking mode' on difficult reports.
Conclusions:
- Current LLMs can extract CAD-RADS stenosis severity categories from unstructured CCTA reports with performance approaching human inter-rater agreement.
- Open-weight models, such as Gemma 4 31B, offer competitive performance and enable privacy-preserving local deployment on consumer hardware.
- LLMs represent a significant advancement for automated clinical data mining and improving reporting consistency in cardiovascular imaging.