Related Experiment Video
Updated: Oct 1, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Real-time Robust Semantic Segmentation via Low-Resolution Ensemble
Abstract:
Semantic segmentation algorithms constitute a critical component for vision-based perception in robotics and autonomous driving. Real-world deployment in uncertain environments requires such algorithms to be both low-latency and robust against common corruptions. However, the limited representation ability of real-time models and the high computational overhead of existing robust models render real-time robust semantic segmentation particularly challenging. In this paper, we begin by systematically evaluating the robustness of existing real-time models and observe that inference on down-sampled images substantially enhances their performance under common corruptions. Further analysis reveals that robustness to high-frequency corruptions plays a pivotal role. Drawing on the Information Bottleneck principle, we provide an interpretation of how image down-sampling influences performance on both clean and corrupted data. Building on these insights, we propose RTRSSeg, an ensemble-based approach dedicated to real-time robust semantic segmentation, which incorporates a robustness branch and a detail branch. The robustness branch integrates a heavyweight vision foundation model that processes down-sampled low-resolution image blocks to enhance feature compression and invariance, while the detail branch employs a lightweight convolutional network to preserve spatial details. To decouple the optimization process, auxiliary supervision is applied exclusively to the robustness branch. Furthermore, we design a multi-scale alignment module to mitigate pattern discrepancies between the two branches and facilitate effective feature fusion. Additionally, a mask refiner is introduced to sharpen predicted masks. Comprehensive experiments and indepth quantitative analyses demonstrate the efficacy of our approach. RTRSSeg achieves notable robustness while operating at 30-90 FPS on a single RTX3090 GPU. The code will be available at https://github.com/ydhongHIT/RTRSSeg.