Related Experiment Video
Updated: Jun 7, 2025

Expedited Radiation Biodosimetry by Automated Dicentric Chromosome Identification ADCI and Dose Estimation
Published on: September 4, 2017
ChatGPT vs Gemini: Comparative Accuracy and Efficiency in CAD-RADS Score Assignment from Radiology Reports
Matthew Silbergleit1, Adrienn Tóth1, Jordan H Chamberlin1
1Division of Cardiothoracic Imaging, Department of Radiology and Radiological Science, Clinical Science Building, Medical University of South Carolina, 96 Jonathan Lucas Street, Suite 210, MSC 323, Charleston, SC, 29425, USA.
ChatGPT-4o achieved the highest accuracy in generating Coronary Artery Disease Reporting and Data System (CAD-RADS) scores from radiology reports. While faster, ChatGPT-3.5 was less accurate, indicating LLMs need further development for clinical use.
Area of Science:
- Artificial Intelligence in Radiology
- Medical Informatics
- Natural Language Processing
Background:
- Radiology reports contain crucial data for disease classification.
- Automating Coronary Artery Disease Reporting and Data System (CAD-RADS) scoring can improve efficiency.
- Large Language Models (LLMs) show potential for analyzing medical text.
Purpose of the Study:
- To evaluate the accuracy and efficiency of ChatGPT-3.5, ChatGPT-4o, Google Gemini, and Google Gemini Advanced in generating CAD-RADS scores.
- To compare LLM performance against radiologist-assigned CAD-RADS scores.
- To assess the interobserver agreement of LLM-generated scores.
Main Methods:
- Retrospective analysis of 100 coronary computed tomography angiography reports.
- Inputting the findings section of reports into four LLMs without fine-tuning.
- Comparing LLM-generated CAD-RADS scores with radiologist-assigned scores.
- Recording the time taken by each LLM.
Main Results:
- ChatGPT-4o demonstrated the highest accuracy (87%) and strong interobserver agreement (κ=0.838, α=0.886).
- Gemini Advanced achieved 82.6% accuracy (κ=0.784, α=0.897).
- ChatGPT-3.5 was the fastest (5s) but least accurate (50.5%, κ=0.401, α=0.787).
- Gemini had a 12% failure rate, with Gemini Advanced showing improvement.
Conclusions:
- ChatGPT-4o is the most accurate LLM for CAD-RADS scoring among those tested.
- Current LLMs require further refinement for reliable clinical decision-making in CAD-RADS scoring.
- LLM speed and accuracy vary significantly, impacting potential clinical integration.
More Related Videos
06:16Signal Acquisition, Score Interpretation, and Economics of a Non-Invasive Point-of-Care Test for Coronary Artery Disease
Published on: August 9, 2024
06:57Author Spotlight: Advancing Cardiovascular Imaging - Introducing the Spatially Weighted Calcium Score for Early Disease Detection
Published on: September 22, 2023