Related Experiment Videos

Artificial Intelligence Performance Under Different Conditions in Answering China's Standardized Training Examination

Zheng Zhu1, Yanfeng Zhao1, Lin Li1

  • 1Department of Diagnostic Radiology National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College Beijing China.

Health Care Science
|July 25, 2026
PubMed
Summary

Large language models (LLMs) show varied performance on radiology exams. Gemini-2.0 achieved the highest accuracy, but introducing doubt did not consistently improve results for these AI models.

Related Concept Videos