Related Experiment Video
Updated: Mar 9, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
3.6K
Multimodal medical endoscopic image analysis via progressive disentangle-aware contrastive learning.
Junhao Wu1, Yun Li2, Junhao Li3
1Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, Guangdong, 518055, China; Department of Computer and Information Sciences, Towson University, Towson, MD, 21252, USA.
Medical Image Analysis
|March 7, 2026
Summary
This study introduces a new AI framework that combines two imaging types, White Light Imaging (WLI) and Narrow Band Imaging (NBI), to improve the segmentation of laryngo-pharyngeal tumors for better diagnosis and treatment.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computer Vision
Background:
- Accurate segmentation of laryngo-pharyngeal tumors is vital for diagnosis and treatment planning.
- Single-modality imaging often lacks the comprehensive features needed for precise tumor delineation.
- Integrating multiple imaging modalities can enhance the understanding of tumor characteristics.
Purpose of the Study:
- To develop and evaluate a multi-modality representation learning framework for improved laryngo-pharyngeal tumor segmentation.
- To address the challenge of modality discrepancies in medical image analysis.
- To enhance the accuracy and robustness of tumor segmentation using combined WLI and NBI data.
Main Methods:
- A novel Align-Disentangle-Fusion framework integrating 2D White Light Imaging (WLI) and Narrow Band Imaging (NBI).
- Multi-scale distribution alignment to harmonize features across different transformer layers and mitigate modality discrepancies.
- Progressive feature disentanglement, including preliminary disentanglement and disentangle-aware contrastive learning, to separate and leverage modality-specific and shared features.
- Multimodal contrastive learning and efficient semantic fusion for enhanced segmentation performance.
Main Results:
- The proposed framework significantly outperforms existing state-of-the-art methods in laryngo-pharyngeal tumor segmentation.
- Demonstrated superior accuracy and robustness across diverse clinical datasets and scenarios.
- Validation of the effectiveness of multi-scale alignment and feature disentanglement strategies.
Conclusions:
- The multi-modality representation learning framework offers a significant advancement in laryngo-pharyngeal tumor segmentation.
- The integration of WLI and NBI through advanced AI techniques improves diagnostic precision.
- The developed method provides a robust and accurate solution for clinical applications, with source code available for further research.

