Related Experiment Video
Updated: Sep 30, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Revisiting Pixel Erasure: A Simple Method for Weakly Supervised Video Instance Segmentation
Abstract:
We aim to develop a simple framework for weakly supervised video instance segmentation (VIS) without pixelwise video annotations that can match or outperform fully supervised counterparts. To this end, we revisit pixel erasure, which is a common data augmentation strategy, and reveal that it can be naturally used to resist label noise within pseudo masks. Thus, we propose PieVIS, which exploits abundant sources of image datasets and requires only box-level annotations for the video dataset. Specifically, PieVIS consists of two stages: pseudo mask generation and self-training with pixel erasure. In the first stage, PieVIS generates pseudo masks for the video data using a model trained with image datasets. Afterward, for self-training, we adopt a simple pixel erasure technique that randomly removes colour information for some regions while keeping the pseudo masks unchanged. For the first time, we show that VIS without video mask annotations can perform comparably to or even outperform state-of-the-art fully supervised methods on YouTube-VIS (YTVIS) 2019 and 2021. More importantly, we further demonstrate that the pixel erasure strategy can also advance other weakly supervised tasks that involve learning from pseudo masks.
