Related Experiment Video
Updated: Jan 13, 2026

Using Light Sheet Fluorescence Microscopy to Image Zebrafish Eye Development
Published on: April 10, 2016
CAFusion: A progressive ConvMixer network for context-aware infrared and visible image fusion
Hafiz Tayyab Mustafa1,2, Hamza Mustafa3, Hassan Alhuzali4
1School of Computer Science and Technology, Zhejiang Normal University, Jinhua, Zhejiang, China.
None:
Image fusion is a challenging task that aims to generate a composite image by combining information from diverse sources. While deep learning (DL) algorithms have achieved promising results, most rely on complex encoders or attention mechanisms, leading to high computational cost and potential information loss during one-step feature fusion. We introduce CAFusion, a DL framework for visible (VI) and infrared (IR) image fusion. In particular, we propose a context-aware ConvMixer block that uniquely integrates dilated convolutions for expanded receptive fields with depthwise separable convolutions for parameter efficiency. Unlike existing CNN or transformer-based modules, our block captures multi-scale contextual information without attention mechanisms, with computational efficiency. Additionally, we employ an attention-based intermodality multi-level progressive fusion strategy, ensuring an adaptive combination of multi-scale modality-specific features. A hierarchical multiscale decoder reconstructs the fused image by aggregating information across different levels, preserving low and high-level details. Comparative evaluations of benchmark datasets demonstrate that CAFusion outperforms recent transformer-based and SOTA DL-based approaches in fusion quality and computational efficiency. In particular, on the TNO benchmark dataset, CAFusion achieves a 0.769 score in the structural similarity index measure, a 2.07 percent increase as compared to the best competing method.
Related Concept Videos
Infrared (IR) Spectroscopy: Overview
Different compounds display unique properties due to their...
IR Frequency Region: Fingerprint Region
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Light Acquisition
Vision
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...

