Related Experiment Video
Updated: Jul 3, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
537
Disentangled Cross-Modal Transformer for RGB-D Salient Object Detection and Beyond
Summary
This study introduces a novel Disentangled Feature Pyramid (DFP) module for RGB-D salient object detection (SOD). DFP reduces fusion ambiguity by disentangling cross-modal contexts and representations, improving performance and adaptability.
Area of Science:
- Computer Vision
- Artificial Intelligence
Background:
- Multi-modal transformers for RGB-D salient object detection (SOD) often fuse features directly, leading to inefficient and ambiguous results due to modality gaps.
- Existing methods struggle to effectively model cross-modal correlations and combine information from different sources without clear differentiation.
Purpose of the Study:
- To propose a novel Disentangled Feature Pyramid (DFP) module to reduce cross-modal fusion ambiguity in RGB-D salient object detection.
- To disentangle cross-modal complementarity into context and representation levels for more informative feature integration.
- To enhance fusion adaptivity and improve the performance of salient object detection models.
Main Methods:
- Context disentanglement: separating long-range complementary contexts for intra-modal self-attention and local correlations via spatial-aligned inter-modal attention.
- Representation disentanglement: dividing tokens into consistent and private channels to explicitly boost complementary information and disentangle integration paths.
- Progressive propagation of the DFP module across layers for enhanced cross-modal, cross-level integration.
Main Results:
- The proposed Disentangled Feature Pyramid (DFP) module significantly reduces cross-modal fusion ambiguity.
- Experiments on multiple public datasets demonstrate consistent improvements over state-of-the-art SOD models.
- The DFP module shows plug-and-play capabilities, enhancing both transformer and CNN backbones for various downstream tasks.
Conclusions:
- Disentangling cross-modal context and representation is crucial for effective RGB-D salient object detection.
- The DFP module offers a flexible and adaptive approach to multi-modal fusion, outperforming existing methods.
- The proposed method generalizes well to different architectures and tasks, including RGB-D semantic segmentation.

