Related Experiment Video
Updated: Jun 13, 2026

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
Published on: October 27, 2023
Patient-Facing Radiology Communication with LLMs: Calibration Deficit and the Metadata Paradox
Cheong Shin1, Jung Hyun Park1,2,3, Sungjun Kim1,3,4,5
1Department of Integrative Medicine, The Graduate School, College of Medicine, Yonsei University, Seoul 03722, Republic of Korea.
Abstract:
Background/Objectives: Patients increasingly access radiology reports via online portals and frequently seek clarification. While Large Language Models (LLMs) may facilitate this communication, their clinical safety and reliability in this context remain largely uncharacterized. This study aimed to evaluate performance heterogeneity (the disparity between factual synthesis and interpretive reasoning), the Metadata Paradox (performance degradation triggered by demographic priors), and calibration characteristics in answering simulated patient questions derived from radiology reports. Methods: In this retrospective study, 2000 simulated inquiries were generated from 200 MIMIC-IV radiology reports based on an expert-refined 10-category taxonomy, categorized into factual tasks (e.g., terminology/anatomy) and interpretive tasks (e.g., diagnostic confidence/finding detail). Three LLMs (GPT-4o mini, Grok (v4-0709), Claude 3.5 Sonnet) generated 12,000 answers (with/without metadata). Quality was scored (1-3 scale) by Gemini 2.5 Flash, validated by three independent board-certified radiologists and finalized through four-specialist consensus adjudication (n = 1200). Performance and self-confidence calibration were assessed using Generalized Estimating Equations. Results: The LLM judge showed an overall agreement rate of 90.5% with the adjudicated ground truth. Grok and Claude 3.5 Sonnet significantly outperformed GPT-4o mini (p < 0.001); specifically, GPT-4o mini was associated with a 2.8-fold higher risk of failure compared to Grok (adjusted OR 2.83; 95% CI: 2.28-3.49; p < 0.001) and an absolute risk difference (ARD) of 8.4 percentage points. Accuracy reached its ceiling in factual tasks (Terminology: 98.1%) but was significantly lower in interpretive tasks (Diagnostic Confidence: 82.3%, p < 0.001). Metadata inclusion triggered the 'Metadata Paradox,' significantly increasing the risk of failure (OR 1.11; p = 0.044). A substantial calibration deficit (defined as the disconnect between self-confidence and accuracy) was observed; notably, the majority of safety-critical errors (Score 1: clinically significant misinformation; n = 131) were assigned high self-confidence (≥8/10; GPT-4o mini: 93.8%, Grok: 100%, Claude 3.5 Sonnet: 61.5%). Conclusions: Although LLMs accurately address factual queries, their consistent calibration deficit in safety-critical errors and susceptibility to stochastic stereotyping highlight the necessity of independent verification frameworks.
Related Concept Videos
Magnetic Resonance Imaging
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
Imaging Studies for Cardiovascular System IV: CMRI
X-ray Imaging
Radiological Investigation I: X-ray and CT
