Related Experiment Video
Updated: May 28, 2026

Lesion Explorer: A Video-guided, Standardized Protocol for Accurate and Reliable MRI-derived Volumetrics in Alzheimer's Disease and Normal Elderly
Published on: April 14, 2014
Performing Best When Needed Least: Reader Experience Shapes Accuracy Gains in Large Language Model-assisted Brain MRI
Severin Schramm1, Bastien Le Guellec2, Marlene Topka3
1Department of Diagnostic and Interventional Neuroradiology, TUM University Hospital, School of Medicine and Health, Technical University of Munich, Munich, Germany.
None:
Background Studies have demonstrated that large language models (LLMs) can perform differential diagnosis based on textual radiologic findings; however, it is unclear how variations in reader-generated inputs affect LLM performance and clinical utility. Purpose To evaluate how reader experience influences the diagnostic benefit of LLM assistance in brain MRI differential diagnosis. Materials and Methods In this retrospective multireader study, neuroradiologists (n = 4), radiology residents (n = 4), and neurology/neurosurgery residents (n = 4) provided textual radiologic findings and their top three differential diagnoses for brain MRI scans with confirmed diagnoses obtained between January 2009 and April 2024 from a single academic center. Confirmed diagnoses were established histopathologically or through consensus of at least two neuroradiologists. Three LLMs (GPT-4.1 [OpenAI], Gemini 2.5 Pro [Google DeepMind], and DeepSeek-R1 [Hangzhou DeepSeek Artificial Intelligence Basic Technology Research]) generated differential diagnoses based on reader-provided findings. Readers revised their diagnoses after reviewing the suggestions of GPT-4.1. A cumulative link mixed model was fitted to evaluate the association between reader experience and diagnostic benefit, with change in diagnostic result as an ordinal outcome, reader experience as a predictor, and random intercepts for rater and patient. Results Forty brain MRI scans (mean patient age, 50 years ± 18 [SD]; 23 female) were included. LLM-generated diagnoses achieved the highest top-three accuracy based on imaging findings from neuroradiologists (78.8%-83.8% across LLMs), followed by radiology residents (71.8%-77.6%) and neurology/neurosurgery residents (63.2%-67.1%). Mean absolute gains in top-three accuracy with LLM assistance diminished with increasing experience: +19.4% for neurology/neurosurgery residents (from 43.2% to 62.6%), +14.7% for radiology residents (from 59.6% to 74.4%), and +4.4% for neuroradiologists (from 83.1% to 87.5%). Models demonstrated a negative association between reader experience and diagnostic benefit from LLM assistance (β = -0.10; P = .005) and a positive association of reader experience with correctness (β = 0.11; P < .001) and completeness (β = 0.18; P = .002) of imaging findings. Conclusion With increasing reader experience, LLM accuracy with reader-generated input improved, whereas accuracy gains from LLM assistance diminished. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by McMillan in this issue.
