Related Experiment Video
Updated: Mar 28, 2026

Multi-modal Pulmonary Imaging: Using Complementary Information from CT and Hyperpolarized 129Xe MRI to Evaluate Lung Structure-Function
Published on: April 12, 2024
Multimodal vision-language models in chest x-ray analysis: a study of generalization, supervision, and robustness
Batoul Aljaddouh1, D Malathi1, Feisal Alaswad1
1Department of Computing Technologies, SRM Institute of Science and Technology, Kattankulathur, Tamil Nadu 603203 India.
None:
Multimodal vision-language models (VLMs) are increasingly applied to medical imaging, yet systematic evaluations comparing them with unimodal models across datasets, supervision regimes, and clinical domains remain scarce. Prior studies often focus on a single dataset, specific pathologies, or one supervision setting, leaving unclear how these models generalize under realistic variability. We conduct a systematic evaluation of six leading unimodal and multimodal models for chest X-ray (CXR) classification using four widely adopted datasets: MIMIC-CXR, CheXpert, NIH-14, and PadChest. We assess model behavior in both zero-shot (ZS) and fine-tuned (FT) configurations, with a focus on generalization across pathologies, datasets, and linguistic domains. Our findings show that pretrained multimodal models such as CheXzero and CXR-LLaVA perform strongly in zero-shot scenarios, especially on out-of-distribution data, reflecting their capacity for semantic generalization. However, their performance tends to decline after fine-tuning in cross-lingual or noisy-label contexts, indicating susceptibility to overfitting. In contrast, unimodal models gain substantially from supervised fine-tuning, especially on in-domain data. Limitations include evaluation on seven shared pathologies, CXR imaging only, and use of publicly available pretrained models, which may restrict generalization to other clinical tasks. These findings highlight key trade-offs between generalization, robustness, and adaptability, and suggest promise in hybrid training strategies that integrate multimodal priors with targeted domain supervision.
