Related Experiment Video
Updated: Sep 20, 2025

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
3.0K
Hybrid Masked Image Modeling for 3D Medical Image Segmentation
IEEE Journal of Biomedical and Health Informatics
|January 30, 2024
Summary
HybridMIM enhances 3D medical image segmentation by using a novel hybrid self-supervised learning approach. This method significantly reduces pre-training time and improves segmentation accuracy compared to existing techniques.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Computer Vision
Background:
- Masked image modeling (MIM) with transformers is a powerful self-supervised pre-training technique.
- Existing MIM methods reconstruct pixels, focusing on low-level semantics and incurring long pre-training times.
- There is a need for more efficient and semantically rich self-supervised methods for 3D medical image segmentation.
Purpose of the Study:
- To introduce HybridMIM, a novel hybrid self-supervised learning method for 3D medical image segmentation.
- To improve the efficiency and effectiveness of self-supervised pre-training for medical image analysis.
- To leverage multi-level semantic information for enhanced segmentation performance.
Main Methods:
- A two-level masking hierarchy is designed to specify masking of patches in sub-volumes, incorporating higher-level semantic constraints.
- Three levels of semantic learning are employed: partial region prediction (pixel-level), patch-masking perception (region-level), and contrastive learning (sample-level).
- The framework supports both Convolutional Neural Networks (CNNs) and transformers as encoder backbones and allows pre-training of decoders.
Main Results:
- HybridMIM demonstrates clear superiority over supervised methods, other masked pre-training approaches, and self-supervised methods on five public medical image segmentation datasets (BraTS2020, BTCV, MSD Liver, MSD Spleen, BraTS2023).
- The method achieves significant improvements in quantitative metrics, speed performance, and qualitative observations.
- Reduced pre-training time was observed due to partial region prediction.
Conclusions:
- HybridMIM offers a versatile and effective self-supervised learning framework for 3D medical image segmentation.
- The proposed multi-level semantic learning strategy enhances segmentation accuracy and generalization ability.
- HybridMIM presents a promising advancement in self-supervised learning for medical image analysis, outperforming existing methods.

