VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing
IEEE Transactions on Bio-Medical Engineering
|August 13, 2026
Summary
VIGIL, a novel vision-language framework, enhances ulcerative colitis histological healing prediction using endoscopy images and reports. This approach improves diagnostic accuracy and reduces annotation needs for better patient outcomes.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Gastroenterology
Background:
- Ulcerative colitis (UC) involves chronic inflammation with cycles of remission and relapse.
- Accurate histological healing (HH) assessment is crucial for improving UC clinical outcomes.
- Current deep learning and multi-instance learning (MIL) methods for HH prediction have limitations.
Purpose of the Study:
- To introduce VIGIL, a pioneering vision-language-guided MIL framework for UC HH prediction.
- To integrate white-light endoscopy (WLE) and endocytoscopy (EC) modalities within the VIGIL framework.
- To overcome the challenges of annotation-intensive methods and improve prediction reliability.
Main Methods:
- VIGIL utilizes a dual-branch MIL module (KS-MIL) with top-K frame selection and similarity-weighted metric learning.
- It incorporates diagnostic report text and a multi-level image-text alignment strategy for joint vision-language guidance.
- A multi-modal masked relation fusion (MMRF) strategy is employed to fuse WLE and EC representations.
Main Results:
- VIGIL achieved 92.69% accuracy and 94.79% AUC on a clinical dataset.
- The framework demonstrated superior performance compared to existing state-of-the-art methods.
- Experiments confirmed the effectiveness of the proposed VIGIL framework.
Conclusions:
- VIGIL establishes an effective vision-language guided MIL paradigm for UC HH prediction.
- The framework reduces annotation burden and enhances prediction reliability.
- This research offers insights for non-invasive UC diagnosis and intelligent healthcare advancement.
