Related Experiment Video
Updated: Aug 8, 2025

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
471
A modality-collaborative convolution and transformer hybrid network for unpaired multi-modal medical image
Hong Liu1, Yuzhou Zhuang1, Enmin Song1
1School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China.
Medical Physics
|March 3, 2023
Summary
This study introduces a new network for segmenting unpaired medical images, effectively handling variations and limited data. The modality-collaborative convolution and transformer hybrid network (MCTHNet) significantly improves segmentation accuracy with fewer labels.
Area of Science:
- Medical image analysis
- Artificial intelligence in healthcare
- Computer vision
Background:
- Multi-modal learning is crucial for medical image segmentation, but traditional methods require paired, aligned images.
- Unpaired multi-modal learning addresses limitations but often overlooks scale variations and struggles with global context.
- Limited labeled data in clinical settings hinders the training of accurate segmentation models.
Purpose of the Study:
- To develop a novel network (MCTHNet) for semi-supervised, unpaired multi-modal medical image segmentation using limited annotations.
- To address intensity distribution gaps, scale variations, and the need for global contextual information.
- To leverage extensive unlabeled scans to enhance segmentation performance, reducing annotation burden.
Main Methods:
- A modality-specific scale-aware convolution (MSSC) module adaptively handles inter-modality differences.
- A modality-invariant vision transformer (MIViT) module captures both local and global features for generalized representations.
- Multi-modal cross pseudo supervision (MCPS) enables semi-supervised learning from unlabeled data.
Main Results:
- The proposed MCTHNet significantly outperforms existing methods on unpaired CT and MR segmentation tasks.
- With only 25% labeled data, the method achieves high segmentation performance (e.g., 78.56% DSC for cardiac).
- Performance approaches that of fully supervised single-modal methods, demonstrating effectiveness with limited annotations.
Conclusions:
- The proposed MCTHNet effectively reduces the annotation burden for multi-modal medical image segmentation.
- This approach facilitates the practical application of advanced segmentation techniques in clinical settings.
- The method demonstrates the potential of semi-supervised learning with hybrid networks for unpaired medical image analysis.

