Related Experiment Video
Updated: Jan 18, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
3.3K
AMVLM: Alignment-Multiplicity Aware Vision-Language Model for Semi-Supervised Medical Image Segmentation.
IEEE Transactions on Medical Imaging
|May 23, 2025
Summary
Low-quality pseudo-labels hinder semi-supervised medical image segmentation. The Alignment-Multiplicity Aware Vision-Language Model (AMVLM) improves pseudo-label quality and segmentation performance by addressing cross-modal alignment challenges.
Area of Science:
- Computer Vision
- Medical Imaging
- Artificial Intelligence
Background:
- Low-quality pseudo-labels are a major challenge in semi-supervised medical image segmentation (SSMIS).
- Vision-Language Models (VLMs) show potential for improving pseudo-label quality but struggle with cross-modal alignment uncertainty.
- Existing VLM approaches can lead to semantic degradation due to distribution-based semantic modeling.
Purpose of the Study:
- To introduce a novel VLM pretraining paradigm, the Alignment-Multiplicity Aware Vision-Language Model (AMVLM).
- To enhance pseudo-label quality and improve consistency learning in SSMIS.
- To develop a text-guided SSMIS network leveraging the pretrained AMVLM.
Main Methods:
- Proposed the Alignment-Multiplicity Aware Vision-Language Model (AMVLM) pretraining paradigm.
- Introduced Cross-modal Similarity Supervision (CSS) using a probability distribution transformer for multi-alignment.
- Implemented Intra-modal Contrastive Learning (ICL) for coarse-fine granularity semantic consistency.
- Developed a text-guided SSMIS network with a text mask generator for multimodal supervision.
Main Results:
- The AMVLM effectively addresses cross-modal alignment uncertainty and semantic degradation.
- The proposed CSS and ICL strategies enable learning of cross-modal multiple alignments and semantic consistency.
- The AMVLM-driven SSMIS network significantly enhances pseudo-label quality and consistency learning.
- Superior performance was demonstrated across four public datasets compared to existing methods.
Conclusions:
- The AMVLM pretraining paradigm offers a robust solution for VLM challenges in SSMIS.
- The developed text-guided SSMIS network effectively leverages multimodal information to improve segmentation accuracy.
- The findings highlight the potential of AMVLM for advancing semi-supervised medical image analysis.
