Related Experiment Video
Updated: Jan 18, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
DMDNet: Dual-branch multi-modal deep fusion network for V-D-T salient object detection
Yaoqi Sun1, Bin Wan2, Haibing Yin3
1School of Artificial Intelligence, Lishui University, Lishui, 323000, China; Lishui Institute of Hangzhou Dianzi University, Hangzhou Dianzi University, Hangzhou, 310018, China.
This study introduces a dual-branch deep fusion network (DMDNet) for multi-modal salient object detection. DMDNet improves accuracy by fusing features later in the process, reducing noise and enhancing detection performance.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Directly fusing multi-modal features (visible, depth, thermal) in early encoding stages introduces noise, reducing salient object detection accuracy.
- Existing methods often struggle with effectively integrating complementary information from different sensor modalities.
Purpose of the Study:
- To propose a novel dual-branch multi-modal deep fusion network (DMDNet) for improved salient object detection.
- To address the challenge of noise amplification in early feature fusion by performing fusion in the decoder phase.
Main Methods:
- DMDNet utilizes separate encoder branches for visible images and a combined branch for depth and thermal images.
- The network incorporates a modal interaction (MI) module for depth-thermal complementarity, multi-scale feature perception (MFP), and region optimization (RO) modules.
- A dual-branch fusion (DF) module integrates features bottom-to-top for final saliency map generation.
Main Results:
- DMDNet demonstrates superior performance on the VDT-2048 dataset.
- Experimental results validate the effectiveness of the proposed network architecture and fusion strategy.
Conclusions:
- The proposed DMDNet effectively reduces noise by delaying multi-modal fusion to the decoder phase.
- The network architecture successfully leverages complementary features from visible, depth, and thermal modalities for enhanced salient object detection.
Related Concept Videos
Depth Perception and Spatial Vision
Uniform Depth Channel Flow
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Multi-input and Multi-variable systems
In the absence of...
Uniform Depth Channel Flow: Problem Solving