Related Experiment Video
Updated: Jan 16, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Enhancing the CAD-RADS™ 2.0 Category Assignment Performance of ChatGPT and DeepSeek Through "Few-shot" Prompting
1Department of Radiology, School of Medicine, Bursa Uludağ University, Bursa, Turkey.
Few-shot prompting significantly improved large language models' performance in assigning Coronary Artery Disease Reporting and Data System (CAD-RADS™ 2.0) categories, achieving high accuracy and eliminating hallucinations.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing for Radiology Reports
- Machine Learning for Clinical Decision Support
Background:
- Large language models (LLMs) show promise in medical text analysis.
- Accurate categorization of Coronary Artery Disease Reporting and Data System (CAD-RADS™ 2.0) is crucial for patient management.
- The performance of LLMs in CAD-RADS™ 2.0 classification using few-shot prompting is not well-established.
Purpose of the Study:
- To evaluate the impact of few-shot prompting on the performance of ChatGPT and DeepSeek LLMs for CAD-RADS™ 2.0 category assignment.
- To compare the accuracy and reproducibility of LLM-based CAD-RADS™ 2.0 classification with radiologist assessments.
Main Methods:
- A few-shot prompt was developed using 20 reports from the MIMIC-IV database based on the CAD-RADS™ 2.0 framework.
- 100 MIMIC-IV reports were classified using zero-shot and few-shot prompts with ChatGPT and DeepSeek.
- Model performance was assessed against a reference radiologist's classifications, with statistical analysis using McNemar tests and Cohen kappa.
Main Results:
- Zero-shot prompting yielded low accuracy (ChatGPT: 14%, DeepSeek: 8%) with frequent hallucinations.
- Few-shot prompting significantly improved accuracy to 98% for ChatGPT and 93% for DeepSeek (P<0.001), eliminating hallucinations.
- High agreement (kappa > 0.91) was observed between LLM-generated and radiologist classifications, with excellent reproducibility.
Conclusions:
- Few-shot prompting substantially enhances LLM accuracy for CAD-RADS™ 2.0 classification.
- This approach demonstrates potential for clinical integration in cardiovascular imaging interpretation.
- LLMs optimized with few-shot prompting can facilitate efficient and accurate adoption of CAD-RADS™ 2.0.
More Related Videos
05:58Using Rapid Serial Visual Presentation to Measure Set-Specific Capture, a Consequence of Distraction While Multitasking
Published on: August 29, 2018
08:17A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018
Related Concept Videos
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Improving Translational Accuracy
Improving Translational Accuracy