Related Experiment Video
Updated: Jan 15, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Multi-modal semi-supervised medical image segmentation via spatial weight fusion and prototype-based alignment
Xiao Tian1, Biyuan Li1,2, Jinying Ma1
1School of Electronic Engineering, Tianjin University of Technology and Education, Tianjin, China, 300222, People's Republic of China.
Abstract:
Multi-modal learning leverages complementary information from different modalities to enhance medical image segmentation. However, existing methods often require large-scale, high-quality annotations, which are scarce in clinical practice, and suffer from anatomical misalignment between modalities.We propose MHPC-Net, a Manhattan Hybrid-attentive Prototype-aligned Cross-modal Network for semi-supervised multi-modal segmentation. MHPC-Net uses a dual-branch design to integrate CT and MRI information, producing anatomically consistent and complementary outputs under limited supervision. A feature interaction module combines spatial weights derived from Manhattan distance with dual-modal cross-attention to enhance inter-modal exchange while preserving fine anatomical details, mitigating modality misalignment. A multi-level feature fusion module achieves deep semantic integration with spatial and anatomical consistency in a lightweight manner. To address semantic inconsistency from modality-specific traits and channel misalignment, a modality contrast strategy projects modality-invariant representations and modality-specific knowledge into distinct spaces. A feature decorrelation loss enforces independence between them, preserving complementary information. A prototype alignment mechanism with a memory bank further refines structure consistency and aligns modality-invariant representations for robust cross-modal representation learning.Experiments on cardiac and abdominal segmentation show that MHPC-Net achieves state-of-the-art performance under limited labels, improving accuracy and generalization in semi-supervised multi-modal scenarios.

