Related Experiment Video
Updated: Sep 19, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
When Grouped Cyclic Shift meets masked image modeling: Effective pre-training for data-scarce 3D ultrasound analysis
Rui Zhou1, Yingtai Li2, Tianzhu Liang3
1School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China (USTC), Hefei, Anhui, 230026, China; Mindray Bio-Medical Electronics Co., Ltd., Shenzhen, China; Medical Imaging, Robotics, Analytic Computing & Learning (MIRACLE) Lab, YRD-RIGHT, USTC Suzhou Institute for Advanced Research, Suzhou, Jiangsu, 215123, China.
Abstract:
The inherent data scarcity in 3D ultrasound analysis demands data-efficient self-supervised learning (SSL) methods, yet prevalent approaches are often data-intensive. To bridge this gap, we present Grouped Cyclic Shift Masked Image Modeling (GCSMIM), a masked image modeling (MIM) pre-training framework designed for this low-data regime. GCSMIM employs a data-efficient architecture, where a convolutional neural network (CNN) backbone operates in shallow, high-resolution layers for hierarchical feature extraction, while a multilayer perceptron (MLP) operates in the deep, low-resolution layers to capture long-range dependencies. This design efficiently enhances feature mixing and captures long-range contextual information, providing structured spatial interaction without self-attention in the deep encoder stages. To further improve the data efficiency of MLPs, we inject human-designed inductive bias with a parameter-free Grouped Cyclic Shift (GCS) operation. A key challenge, however, is that naively applying shifts within MIM causes mask-feature misalignment. We solve this with a novel Mask-Guided Feature Reformation (MGFR) mechanism, which synchronously shifts both the features and the mask, then selectively reintegrates features from the original spatial context, thereby preserving consistency between shifted features and mask states. A Sparse MLP (SMLP) further processes features at mask-identified valid positions during pre-training. Pre-trained on a large-scale dataset of over 1000 3D ultrasound volumes, GCSMIM achieves its clearest gains in the evaluated lower-label setting. Under the stated same-hardware protocols, it also has a shorter fine-tuning step time, lower inference latency, and lower peak allocated memory than a closely related hierarchical CNN-ViT MIM baseline. These findings support label-efficient transfer, particularly under limited annotation budgets. Code is publicly available at https://github.com/MohuaChou/GCSMIM.git.
More Related Videos
09:10Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
16:01An Experimental Protocol for Assessing the Performance of New Ultrasound Probes Based on CMUT Technology in Application to Brain Imaging
Published on: September 24, 2017