Related Experiment Video
Updated: Jun 13, 2026

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
Published on: October 27, 2023
Patient-Facing Radiology Communication with LLMs: Calibration Deficit and the Metadata Paradox
Cheong Shin1, Jung Hyun Park1,2,3, Sungjun Kim1,3,4,5
1Department of Integrative Medicine, The Graduate School, College of Medicine, Yonsei University, Seoul 03722, Republic of Korea.
Large Language Models (LLMs) show promise in answering patient questions about radiology reports, but exhibit significant calibration deficits and biases. Independent verification is crucial for ensuring clinical safety and reliability of LLM-generated responses.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Radiology Patient Communication
Background:
- Patients increasingly access radiology reports online and require clarification.
- Large Language Models (LLMs) may aid patient comprehension but their clinical safety is uncharacterized.
- This study evaluates LLM performance in answering patient questions derived from radiology reports.
Purpose of the Study:
- To assess LLM performance heterogeneity, including factual synthesis vs. interpretive reasoning.
- To investigate the 'Metadata Paradox' where demographic priors degrade LLM performance.
- To evaluate LLM calibration characteristics in answering simulated patient queries.
Main Methods:
- Generated 2000 simulated patient inquiries from 200 MIMIC-IV radiology reports.
- Categorized inquiries into factual (e.g., terminology) and interpretive (e.g., diagnostic confidence) tasks.
- Evaluated three LLMs (GPT-4o mini, Grok, Claude 3.5 Sonnet) on 12,000 generated answers using expert-adjudicated quality scoring.
Main Results:
- Grok and Claude 3.5 Sonnet outperformed GPT-4o mini; GPT-4o mini had a 2.8-fold higher failure risk.
- LLMs excelled in factual tasks (Terminology: 98.1%) but struggled with interpretive tasks (Diagnostic Confidence: 82.3%).
- The 'Metadata Paradox' increased failure risk; significant calibration deficits were observed, with high confidence in safety-critical errors.
Conclusions:
- LLMs accurately answer factual radiology report queries but exhibit calibration deficits in interpretive tasks.
- The 'Metadata Paradox' and high confidence in misinformation necessitate caution.
- Independent verification frameworks are essential before deploying LLMs for patient communication regarding radiology reports.
Related Concept Videos
Magnetic Resonance Imaging
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
Imaging Studies for Cardiovascular System IV: CMRI
X-ray Imaging
Radiological Investigation I: X-ray and CT
