Related Experiment Videos
Artificial Intelligence Performance Under Different Conditions in Answering China's Standardized Training Examination
Zheng Zhu1, Yanfeng Zhao1, Lin Li1
1Department of Diagnostic Radiology National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College Beijing China.
Health Care Science
|July 25, 2026
Summary
Large language models (LLMs) show varied performance on radiology exams. Gemini-2.0 achieved the highest accuracy, but introducing doubt did not consistently improve results for these AI models.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Radiology Training
Background:
- Systematic comparison of general-purpose large language models (LLMs) on specialized medical examinations is lacking.
- Evaluating LLM performance on the China Standardized Training Examination for Resident Physicians in Radiology is crucial.
Purpose of the Study:
- To compare the performance of DeepSeek-R1, ChatGPT-o1, and Gemini-2.0 on radiology resident exams.
- To assess how question formats (with/without choices) and doubt introduction affect LLM accuracy.
Main Methods:
- 131 radiology exam questions were analyzed.
- LLMs were tested on questions with and without answer choices, under conditions of no doubt, weak doubt, and strong doubt.
- Subjective evaluations used a 5-point Likert scale.
Main Results:
- Gemini-2.0 demonstrated the highest accuracy (0.763-0.809).
- LLM accuracy decreased when answer choices were absent, with DeepSeek-R1 showing a statistically significant drop.
- Introducing doubt did not consistently improve accuracy across models and conditions.
- LLMs performed better on single-choice and case-based questions.
Conclusions:
- LLM proficiency varies significantly for radiology resident examinations, with Gemini-2.0 leading in accuracy.
- Repeated questioning or introducing doubt does not reliably enhance LLM performance on radiology-related queries.