Related Experiment Video
Updated: Jan 17, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Hi-End-MAE: Hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image
Fenghe Tang1, Qingsong Yao2, Wenxin Ma1
1School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China (USTC), Hefei, Anhui, 230026, PR China; Center for Medical Imaging, Robotics, and Analytic Computing & LEarning (MIRACLE), Suzhou Institute for Advanced Research, USTC, Suzhou 215123, China; Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Jiangsu, 215123, China.
This study introduces Hierarchical Encoder-driven MAE (Hi-End-MAE), a novel Vision Transformer (ViT) pre-training method. Hi-End-MAE enhances medical image segmentation by better utilizing multi-layer representations for improved accuracy.
Area of Science:
- Medical Imaging
- Computer Vision
- Artificial Intelligence
Background:
- Medical image segmentation faces challenges due to limited labeled data.
- Vision Transformer (ViT) pre-training with masked image modeling (MIM) offers computational efficiency and generalization.
- Existing ViT-based MIM methods often overlook rich, multi-layer representations crucial for fine-grained medical image analysis.
Purpose of the Study:
- To address the limitations of current ViT-based MIM pre-training frameworks in medical imaging.
- To introduce a novel pre-training solution, Hierarchical Encoder-driven MAE (Hi-End-MAE), for improved medical image segmentation.
- To enhance the exploitation of multi-layer representations within ViTs for precise medical downstream tasks.
Main Methods:
- Developed Hierarchical Encoder-driven MAE (Hi-End-MAE), a ViT-based pre-training method.
- Incorporated two key innovations: encoder-driven reconstruction and hierarchical dense decoding.
- Pre-trained Hi-End-MAE on a large dataset of 10,000 CT scans.
Main Results:
- Hi-End-MAE demonstrated superior transfer learning capabilities across nine medical image segmentation benchmarks.
- The method effectively captures fine-grained semantic information by leveraging hierarchical representations.
- Achieved state-of-the-art or competitive performance in various downstream medical imaging tasks.
Conclusions:
- Hi-End-MAE offers a simple yet effective solution for ViT-based pre-training in medical imaging.
- The proposed approach highlights the potential of ViTs for advancing medical image analysis and segmentation.
- Encoder-driven reconstruction and hierarchical decoding are crucial for learning informative features in medical image segmentation.
