Related Experiment Video
Updated: Jan 9, 2026

07:13
Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
1.6K
MedFMIT: A Foundation Model-Driven Multimodal Fusion Method for Image-Text Disease Diagnosis
Summary
A new foundation model, MedFMIT, enhances disease diagnosis by fusing medical images and text. This multimodal approach improves accuracy over image-only methods, offering more precise diagnostic potential.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Analysis
- Natural Language Processing for Healthcare
Background:
- Multimodal disease diagnostic models integrating image and text offer improved precision over image-only approaches.
- Current multimodal methods face challenges with image-text distribution discrepancies and limited cross-modal information interaction.
- Addressing these limitations is crucial for advancing AI-driven medical diagnosis.
Purpose of the Study:
- To propose MedFMIT, a novel foundation model-driven multimodal fusion model for enhanced image-text disease diagnosis.
- To effectively reduce the distributional gap between medical images and text using DCA encoders.
- To enable comprehensive coarse-grained and fine-grained information interaction between visual and textual features via the DMII module.
Main Methods:
- Developed MedFMIT, a foundation model-driven multimodal fusion architecture.
- Employed DCA encoders for aligned visual and textual feature extraction, minimizing distribution differences.
- Integrated a Dual Modality Information Interaction (DMII) module for deep feature fusion.
- Evaluated performance on two multimodal medical classification datasets (CT and endoscopic images).
Main Results:
- MedFMIT achieved superior performance compared to state-of-the-art multimodal algorithms.
- Demonstrated high diagnostic accuracy with Area Under the Curve (AUC) scores of 92.7% and 85.4% on the evaluated datasets.
- Effectively reduced the image-text distributional gap and facilitated robust information interaction.
Conclusions:
- MedFMIT shows significant potential for precise medical diagnosis by effectively fusing image and text data.
- The proposed model architecture addresses key limitations in existing multimodal diagnostic systems.
- Foundation model-driven approaches offer a promising direction for advancing multimodal medical AI.

