Related Experiment Videos
MNet: A Mamba-based teacher-student-decoder framework for multimodal remote sensing segmentation
Mengmeng Liu1, Ronghua Shang1, Jingyu Zhong1
1School of Artificial Intelligence, Xidian University, Xi'an, Shaanxi, 710071, China.
Abstract:
Remote sensing image segmentation serves as a fundamental step in connecting raw data with high-level vision understanding. However, single-modality inputs, such as optical or SAR imagery, often provide incomplete semantics, thereby limiting segmentation accuracy. While recent state space models (e.g., Mamba) have shown strong performance in unimodal segmentation by modeling long-range dependencies efficiently, their potential in multimodal settings remains largely unexplored. To address this gap, we propose MNet, a Mamba-based Teacher-Student-Decoder framework for multimodal remote sensing segmentation. First, we develop a self-adaptive weighting fusion strategy that employs attention-guided, learnable weighting to dynamically integrate optical and SAR features, enhancing cross-modal complementarity, suppressing redundancy, and producing a unified multimodal representation. Second, we introduce an omnidirectional additive state space block, which leverages eight-directional selective scanning and residual aggregation to capture anisotropic spatial features from an overhead perspective. Finally, a teacher network provides global semantic supervision, guiding the student network toward consistent local-global feature alignment, while a decoder refines structural details for precise segmentation. Extensive experiments on the WHU-OPT-SAR and DDHR-Pohang datasets demonstrate that MNet achieves state-of-the-art performance on both multimodal and unimodal tasks, highlighting the potential of Mamba-based architectures for multimodal remote sensing analysis.