Related Experiment Video
Updated: May 9, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Comparative Performance of Seven Mainstream Large Language Models on the 2022 American College of Radiology
Kian A Huang1, David Samvelian1, Allan L Xu1
1Radiology, University of South Florida Morsani College of Medicine, Tampa, USA.
Contemporary large language models (LLMs) show moderate accuracy on radiology exams, outperforming previous models but struggling with image interpretation. While advanced LLMs show promise, they are not yet suitable for diagnostic reasoning tasks requiring visual analysis.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging and Radiology
- Natural Language Processing
Background:
- Large language models (LLMs) show potential in medical education, but comparisons on radiology-specific assessments are limited.
- Recent LLMs like Grok 4.1, Bing Copilot GPT-5, DeepSeek V3.2, and OpenEvidence have not been evaluated on the ACR DXIT examination.
Purpose of the Study:
- To compare the performance of seven contemporary LLMs on the 2022 American College of Radiology Diagnostic Imaging In-Training (ACR DXIT) examination.
- To stratify LLM performance by question format (written-only vs. image-based) and radiology subject domain.
Main Methods:
- Seven LLMs were tested on 106 multiple-choice questions from the 2022 ACR DXIT exam.
- Five multimodal models and two text-only models were used, with a standardized prompt.
- Statistical analysis included Cochran's Q test and McNemar's test, with 95% confidence intervals calculated.
Main Results:
- Multimodal LLMs achieved 65.1%-76.4% accuracy, with no significant differences between models.
- All multimodal models performed significantly better on written (88.1%-95.2%) than image-based (46.9%-64.1%) questions.
- Strengths were noted in ultrasound and chest radiology, while musculoskeletal imaging showed weakness.
Conclusions:
- Contemporary multimodal LLMs demonstrate moderate accuracy on radiology in-training exams, surpassing previous benchmarks.
- A persistent performance gap exists between written and image-based questions, indicating limitations in radiologic image interpretation.
- Current LLMs may aid radiology education for non-interpretive content but are unsuitable for visual diagnostic reasoning.
Related Concept Videos
Imaging Studies III: Computed Tomography
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Imaging Studies VII: Vascular Imaging
Imaging Studies IV: Magnetic Resonance Imaging
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...