Related Experiment Video
Updated: Jan 15, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
3.3K
EviVLM: When Evidential Learning Meets Vision Language Model for Medical Image Segmentation
IEEE Transactions on Medical Imaging
|October 16, 2025
Summary
This study introduces the Evidence-driven Vision Language Model (EviVLM) to bridge the modality gap in medical image segmentation. EviVLM enhances multi-modal fusion for improved segmentation performance using evidential learning.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Medical Imaging
Background:
- The modality gap between image and text representations hinders Vision Language Model (VLM) performance in medical image segmentation.
- This gap complicates multi-modal fusion, limiting segmentation accuracy.
Purpose of the Study:
- To propose a novel paradigm, Evidence-driven Vision Language Model (EviVLM), to systematically measure and mitigate the modality gap.
- To enhance multi-modal fusion for improved medical image segmentation.
Main Methods:
- Integration of Evidential Learning (EL) into VLMs.
- Development of an Evidence Affinity Map Generator (EAMG) for cross-modal evidence collection and refinement.
- Implementation of Evidence Differential Similarity Learning (EDSL) using Bias-Variance Decomposition for consistent cross-modal evidence.
- Application of subjective logic and Dempster-Shafer's theory for evidence mapping, opinion aggregation, and modality gap quantification.
Main Results:
- The proposed EviVLM achieves state-of-the-art performance on three public medical image segmentation datasets.
- Systematic measurement and mitigation of the modality gap were demonstrated.
- Effective multi-modal integration was facilitated, leading to enhanced segmentation.
Conclusions:
- EviVLM offers a novel and effective approach to address the modality gap in medical image segmentation.
- The integration of evidential learning significantly improves multi-modal fusion and segmentation performance.
- The proposed methods provide a robust framework for quantifying and bridging the modality gap in VLMs.

