Related Experiment Video
Updated: Feb 28, 2026

Modeling Brain Metastases Through Intracranial Injection and Magnetic Resonance Imaging
Published on: June 7, 2020
Impact of Large Language Model Assistance on Radiologists' Diagnostic Performance for Brain Tumors by Experience
Chae Won Song1, Byung Hyun Baek1,2, Seul Kee Kim2,3
1Department of Radiology, Chonnam National University Hospital, Gwangju 61469, Republic of Korea.
Large language models (LLMs) like ChatGPT-4o and Claude 3.5 Sonnet show promise in assisting with brain tumor MRI interpretation. LLM assistance significantly improved diagnostic accuracy for radiology trainees, highlighting their potential in clinical settings.
Area of Science:
- Artificial Intelligence in Radiology
- Medical Imaging Analysis
- Machine Learning in Healthcare
Background:
- Large language models (LLMs) offer potential tools for enhancing diagnostic capabilities in medical imaging.
- The diagnostic accuracy of LLMs in interpreting complex cases like brain tumors requires thorough evaluation.
- Comparing LLM performance against human experts is crucial for understanding their role in clinical workflows.
Purpose of the Study:
- To compare the diagnostic accuracy of ChatGPT-4o and Claude 3.5 Sonnet against board-certified radiologists and trainees in brain tumor MRI interpretation.
- To assess whether LLM assistance can improve the diagnostic performance of human readers.
- To evaluate the differential diagnostic capabilities of LLMs based on structured imaging reports.
Main Methods:
- 127 histologically confirmed brain tumor cases were analyzed.
- Two LLMs processed MRI images with structured reports; radiologists and trainees reviewed images with basic demographics.
- Differential diagnoses were generated, and accuracy was calculated before and after LLM assistance.
Main Results:
- Claude 3.5 Sonnet (50.4% primary, 85.0% top-three) and ChatGPT-4o (44.9% primary, 82.7% top-three) showed comparable LLM performance.
- Radiologists outperformed LLMs in primary diagnosis (69.3%) but matched in top-three differentials (80.7%).
- LLM assistance significantly improved top-three differential accuracy for radiologists (80.7% to 90.2%) and both primary (48.0% to 58.8%) and top-three (62.5% to 81.1%) accuracy for trainees.
Conclusions:
- LLMs can expand differential diagnostic considerations when provided with structured imaging data.
- LLM assistance demonstrated a significant positive impact on diagnostic performance, particularly for radiology trainees.
- Further validation is needed under more balanced input conditions and realistic clinical workflows.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
12:50Lesion Explorer: A Video-guided, Standardized Protocol for Accurate and Reliable MRI-derived Volumetrics in Alzheimer's Disease and Normal Elderly
Published on: April 14, 2014