A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology

Qingxia Wu1, Qingxia Wu2, Peipei Zhang3

  • 1Department of Medical Imaging, Henan Provincial People's Hospital & People's Hospital of Zhengzhou University, No.7 Weiwu Road, Zhengzhou, Henan, 450003, China, 86 037165580267, 86 037165651056.

Summary

Multimodal large language models show improved diagnostic performance with images but struggle with 3D data and real-world clinical cases. Bilingual radiology benchmarks reveal performance gaps, indicating models are not yet clinically actionable.

Related Concept Videos