Related Experiment Video
Updated: Sep 11, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
681
One Year On: Assessing Progress of Multimodal Large Language Model Performance on RSNA 2024 Case of the Day Questions
Benjamin Hou1, Pritam Mukherjee2, Vivek Batheja2
1Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, Md.
Radiology
|August 12, 2025
Summary
Recent advancements in multimodal large language models (LLMs) show significant progress in interpreting radiology cases. New OpenAI models demonstrate performance comparable to expert radiologists.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Radiology AI
- Multimodal Large Language Models
Background:
- Growing adoption of multimodal large language models (LLMs) with vision capabilities.
- Need to quantify the performance of these AI models in radiology.
Purpose of the Study:
- Assess and quantify advancements in multimodal LLMs for interpreting radiologic quiz cases.
- Compare AI model performance against senior radiologists.
Main Methods:
- Retrospective study using 95 RSNA 2024 'Case of the Day' questions.
- Baseline comparison with 76 questions from RSNA 2023.
- Evaluated OpenAI, Google Gemini, and Meta Llama models.
- Compared AI accuracy with two senior radiologists using McNemar test.
Main Results:
- OpenAI o1 (59%) and GPT-4o (54%) showed high accuracy on 2024 cases.
- Gemini 1.5 Pro (36%) and Llama 3.2 (33%) had lower scores.
- OpenAI o1 accuracy was comparable to radiologists (58% and 66%).
Conclusions:
- Multimodal LLMs show substantial advancements in one year.
- OpenAI's latest models outperform Google and Meta models.
- No statistically significant difference found between OpenAI o1 and radiologist accuracy.

