Related Experiment Video
Updated: May 24, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
451
A dual-stream feature decomposition network with weight transformation for multi-modality image fusion
Tianqing Hu1, Xiaofei Nan1, Xiabing Zhou2
1School of Computer Science and Artificial Intelligence, Zhengzhou University, Zhengzhou, 450001, China.
Scientific Reports
|March 3, 2025
Summary
This study introduces a novel multi-modal image fusion model combining Transformers and CNNs to enhance infrared and medical imaging. The method effectively fuses features, preserving details for improved object detection and visual tasks.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Image Processing
Background:
- Multi-modal image fusion aims to integrate information from diverse sources into a single image.
- Existing methods using Convolutional Neural Networks (CNNs) have limited receptive fields, while Transformer-based methods are computationally intensive.
- Current approaches often fail to adequately explore cross-domain information and handle feature extraction across different frequencies.
Purpose of the Study:
- To propose an innovative image fusion model for multi-modal images, including infrared/visible and medical imaging pairs.
- To leverage the complementary strengths of Transformers and CNNs for effective modeling of various feature types and learning ranges.
- To enhance downstream visual tasks by generating fused images with preserved salient information and complementary features.
Main Methods:
- A shared Transformer-based encoder for long-range learning, incorporating intra-modal, inter-modal, and feature alignment blocks.
- A private CNN-based dual-stream encoder for low- and high-frequency feature extraction, featuring a dual-domain selection mechanism and an invertible neural network.
- A cross-attention-based Swin Transformer block with embedded weight transformation for efficient cross-domain information exploration and a unified loss function with dynamic weighting.
Main Results:
- The proposed model effectively preserves thermal targets and background texture details in fused images.
- Qualitative and quantitative analyses demonstrate superior performance compared to state-of-the-art methods in image fusion.
- Significant improvements were observed in the performance of subsequent visual tasks, such as object detection.
Conclusions:
- The developed multi-modal image fusion model successfully integrates features from diverse sources using a hybrid Transformer-CNN architecture.
- The model addresses limitations of existing methods by enhancing feature extraction, cross-domain information exploration, and efficiency.
- The proposed approach offers a robust solution for high-quality image fusion, benefiting various downstream computer vision applications.
Related Concept Videos
Deconvolution
127
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
127
Multi-input and Multi-variable systems
93
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
93
Weighted Mean
4.9K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
4.9K
Convolution Properties II
166
The important convolution properties include width, area, differentiation, and integration properties.
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
166
Convolution Properties I
131
Convolution computations can be simplified by utilizing their inherent properties.
The commutative property reveals that the input and the impulse response of an LTI (Linear Time-Invariant) system can be interchanged without affecting the output:
The commutative property reveals that the input and the impulse response of an LTI (Linear Time-Invariant) system can be interchanged without affecting the output:
131
Reducing Line Loss
141
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
141

