Related Experiment Video
Updated: Jan 9, 2026

Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
MedFMIT: A Foundation Model-Driven Multimodal Fusion Method for Image-Text Disease Diagnosis
Abstract:
Image-text multimodal disease diagnostic models have the potential to provide more precise diagnosis results compared to conventional image-only diagnostic models. Existing image-text multimodal diagnostic methods struggle to address the significant image-text distribution differences and realize comprehensive information interaction between the two modalities. To tackle these issues, we propose a novel foundation model-driven multimodal fusion model, MedFMIT, for image-text disease diagnosis. MedFMIT utilizes the DCA encoders to extract informative and well-aligned visual and textual feature representations, effectively reducing the distributional gap between images and texts while ensuring robust feature extraction. The DMII module is introduced to facilitate comprehensive information interaction between image and text features at both coarse-grained and fine-grained levels. For performance evaluation, we conducted experiments on two multimodal medical classification datasets (image + text) containing computed tomography images and endoscopic optical images. MedFMIT outperforms other state-of-the-art multimodal algorithms, achieving the AUC scores of 92.7% and 85.4% on the two datasets respectively, demonstrating its strong potential for precise medical diagnosis.

