Related Experiment Video
Updated: Jun 6, 2025

14:13
Coordinate Mapping of Hyolaryngeal Mechanics in Swallowing
Published on: May 6, 2014
18.2K
Large language models in radiology: Fluctuating performance and decreasing discordance over time
Mitul Gupta1, John Virostko2, Christopher Kaufmann1
1The University of Texas at Austin, Dell Medical School, Department of Diagnostic Medicine, Austin, TX, USA.
European Journal of Radiology
|November 24, 2024
Summary
This study evaluated large language models (LLMs) in radiology, finding GPT-4 initially most accurate but declining over time. Continuous benchmarking is crucial for assessing LLM reliability in medical applications.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Informatics
- Natural Language Processing in Healthcare
Background:
- Large Language Models (LLMs) demonstrate near-expert performance in medical specialties like radiology.
- Limited comparative data exists on LLM performance, accuracy, and reliability over time in radiology.
Purpose of the Study:
- To evaluate and monitor the performance and internal reliability of LLMs in radiology over a three-month period.
- To establish a foundational benchmark for future LLM performance evaluations in radiology.
Main Methods:
- Four LLMs (GPT-4, GPT-3.5, Claude, Google Bard) were queried monthly from November 2023 to January 2024.
- ACR Diagnostic in Training Exam (DXIT) practice questions were used to assess model accuracy by subspecialty.
- Internal consistency was evaluated via answer mismatch or intra-model discordance.
Main Results:
- GPT-4 achieved the highest initial accuracy (78%), followed by Bard (73%), Claude (71%), and GPT-3.5 (63%).
- GPT-4's accuracy trended downwards (82% to 74%), while Claude's increased (70% to 73%) over the study period.
- Intra-model discordance decreased for all LLMs, indicating improved consistency; performance varied by subspecialty.
Conclusions:
- LLMs (excluding GPT-3.5) exceeded 70% accuracy, showing significant domain knowledge.
- Fluctuating performance highlights the need for continuous, radiology-specific benchmarking metrics.
- Standardized evaluation is essential to gauge LLM reliability prior to clinical integration.

