Related Experiment Video
Updated: Jul 5, 2026

13:01
Using Light Sheet Fluorescence Microscopy to Image Zebrafish Eye Development
Published on: April 10, 2016
34.8K
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion.
Summary
This study introduces the Interactive Spatial-Frequency Fusion Mamba (ISFM) for multi-modal image fusion. ISFM enhances image fusion by interactively integrating spatial and frequency domain information, outperforming existing methods.
Area of Science:
- Computer Vision
- Image Processing
- Artificial Intelligence
Background:
- Multi-Modal Image Fusion (MMIF) combines images from diverse sources to preserve details and information.
- Existing methods often use basic spatial-frequency fusion without interactive enhancement.
- Incorporating frequency domain information can improve spatial feature representation in MMIF.
Purpose of the Study:
- To propose a novel Interactive Spatial-Frequency Fusion Mamba (ISFM) framework for advanced MMIF.
- To enhance feature extraction and fusion by integrating spatial and frequency domain information interactively.
- To improve the performance of MMIF systems through a novel fusion strategy.
Main Methods:
- Developed a Modality-Specific Extractor (MSE) for efficient, long-range feature extraction.
- Introduced Multi-scale Frequency Fusion (MFF) for adaptive integration of frequency components.
- Proposed an Interactive Spatial-Frequency Fusion (ISF) module to guide spatial features with frequency information.
Main Results:
- The ISFM framework demonstrated superior performance across six MMIF datasets.
- Experimental results confirmed the effectiveness of the proposed interactive fusion approach.
- ISFM achieved better results compared to current state-of-the-art MMIF methods.
Conclusions:
- The proposed ISFM framework offers a significant advancement in Multi-Modal Image Fusion.
- Interactive integration of spatial and frequency information is crucial for enhanced fusion performance.
- ISFM provides a robust and effective solution for combining multi-modal image data.

