Related Experiment Video
Updated: Oct 10, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
TEAM: Text-Augmented Medical Image Segmentation with Entity-Level Modal Alignment and Masked Modeling
Abstract:
Text-augmented medical image segmentation has become a promising approach to improving medical image analysis by incorporating linguistic descriptions into dense prediction, especially for identifying and segmenting complex anatomical structures. However, existing methods often face challenges due to the semantic gap between concise textual descriptions and dense visual context. This would easily cause the over-alignment or under-alignment issues, resulting in suboptimal segmentation performance. In this paper, we propose a novel framework named TEAM to improve the alignment between visual and textual features and mitigate the potential over-alignment and under-alignment problems. Specifically, our approach introduces an entity contrastive learning (ECL) strategy to enhance the fine-grained correspondence between textual and visual features at their semantic entity levels. In addition, we develop the vision-language feature reconstruction (VL-R) module, which combines masked vision modeling (MVM) with masked latent language modeling (MLM) to further strengthen multimodal feature alignment while preserving the critical and discriminative visual features for precise segmentation. Evaluated on the QaTa-COV19, MosMedData+, and Kvasir-SEG benchmarks, our framework consistently outperforms state-of-the-art baselines. Extensive ablation studies and qualitative visualizations further confirm the effectiveness of our design in enhancing cross-modal alignment and segmentation accuracy, highlighting its potential as a robust solution for text-augmented medical image segmentation.