Large language models in radiology: Fluctuating performance and decreasing discordance over time

Mitul Gupta1, John Virostko2, Christopher Kaufmann1

  • 1The University of Texas at Austin, Dell Medical School, Department of Diagnostic Medicine, Austin, TX, USA.

PubMed
Summary

This study evaluated large language models (LLMs) in radiology, finding GPT-4 initially most accurate but declining over time. Continuous benchmarking is crucial for assessing LLM reliability in medical applications.

Related Concept Videos