Related Experiment Video
Updated: Aug 9, 2026

07:13
Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
Published on: October 27, 2023
A Bilingual Benchmark for Evaluating Diagnostic Performance of Multimodal Large Language Models in Radiology
Qingxia Wu1, Qingxia Wu2, Peipei Zhang3
1Department of Medical Imaging, Henan Provincial People's Hospital & People's Hospital of Zhengzhou University, No.7 Weiwu Road, Zhengzhou, Henan, 450003, China, 86 037165580267, 86 037165651056.
Journal of Medical Internet Research
|August 7, 2026
Summary
Multimodal large language models show improved diagnostic performance with images but struggle with 3D data and real-world clinical cases. Bilingual radiology benchmarks reveal performance gaps, indicating models are not yet clinically actionable.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Analysis
- Natural Language Processing in Healthcare
Background:
- Multimodal large language models (LLMs) are increasingly applied to radiological diagnosis.
- Systematic evaluation of LLM performance across 3D imaging, clinical settings, and bilingual contexts is lacking.
Purpose of the Study:
- To develop a bilingual radiology benchmark (RadM-Bench) for evaluating LLMs.
- To assess LLM diagnostic performance across input modalities, clinical settings, disease rarity, and language.
- To differentiate linguistic effects from clinical content effects via cross-linguistic experiments.
Main Methods:
- Constructed RadM-Bench with 720 cases (360 English teaching, 360 Chinese clinical).
- Evaluated 10 LLMs (4 proprietary, 6 open-source) using history alone, history with 2D images, and history with 3D volumetric data (2 & 10 fps).
- Scored responses using a 4-tier rubric by blinded radiologists; performed cross-lingual translation analysis.
Main Results:
- Mean performance across models was below 1.5/3.0; 2D images improved performance (+19.8% to +139.2%).
- Volumetric data (10 fps) underperformed 2D images (-5.3% to -31.4%); some models had constraints with 3D data.
- Proprietary models declined from teaching to clinical datasets, while Chinese open-source models improved; English translation of Chinese histories decreased scores.
Conclusions:
- Multimodal inputs enhance LLM performance over text alone, but significant gaps exist in 3D data processing and generalization.
- Current LLM diagnostic performance, even with multimodal data, remains below clinically actionable levels.
- Further research is needed to improve LLM capabilities for real-world clinical radiology, especially with volumetric data and diverse linguistic contexts.