Related Experiment Videos
M3amba: Multi-Modality Mamba for Image Fusion
Abstract:
Mamba, a global context modeling paradigm with a selective scanning mechanism, has recently attracted increasing attention in multimodal image fusion. Multimodal fusion aims to preserve and enhance cross-modal complementary information to produce high-quality fused images. Despite its suitability, existing Mamba-based fusion methods often blur modality-specific feature boundaries, mine complementary cues insufficiently, and lack explicit modal-specific modeling. To address these limitations, inspired by cross-attention, we extend Mamba to a multimodal setting and propose a Cross-Mamba architecture. As a plug-and-play module, Cross-Mamba adaptively mines inter-modal complementary information via cross-modal screening, thereby strengthening cross-modal interactions. Moreover, to emphasize complementary cues, reduce redundancy, and lower aggregation complexity, we impose rank constraints to preserve the spatial low-rank structure of features, retaining more informative components and improving fusion quality. Finally, we introduce a masked semantic guidance mechanism to narrow the semantic gap between fusion and downstream tasks, enhancing downstream adaptability. Extensive experiments on multiple fusion datasets, including qualitative, quantitative, and ablation studies, demonstrate that our method achieves state-of-the-art performance.