Related Experiment Video
Updated: Jul 11, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
564
Swin Transformer-Based Edge Guidance Network for RGB-D Salient Object Detection
Shuaihui Wang1, Fengyi Jiang1, Boqian Xu1
1Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences, Changchun 130033, China.
Sensors (Basel, Switzerland)
|November 14, 2023
Summary
This study introduces SwinEGNet, a novel network for salient object detection (SOD) using RGB-D data. SwinEGNet leverages the Swin Transformer to capture global context, outperforming existing methods with enhanced feature fusion and edge guidance.
Area of Science:
- Computer Vision
- Deep Learning
- Image Processing
Background:
- Salient Object Detection (SOD) is crucial for computer vision.
- Existing RGB-D SOD methods using Convolutional Neural Networks (CNNs) suffer from limited performance due to inherent locality.
- There is a need for methods that can capture global context and effectively fuse multi-modal features.
Purpose of the Study:
- To propose a novel Swin Transformer-based edge guidance network (SwinEGNet) for RGB-D SOD.
- To address the limitations of CNN-based methods by incorporating global context extraction.
- To enhance feature fusion through an edge-guided cross-modal interaction module.
Main Methods:
- Employed Swin Transformer as a backbone for feature extraction from RGB and depth images.
- Introduced an Edge Extraction Module (EEM) and a Depth Enhancement Module (DEM).
- Utilized a Cross-Modal Interaction Module (CIM) for integrating global and local cross-modal features, followed by a cascaded decoder for refinement.
Main Results:
- SwinEGNet achieved state-of-the-art performance on LFSD, NLPR, DES, and NJU2K datasets.
- Demonstrated comparable performance on the STEREO dataset against 14 other methods.
- Achieved superior performance over SwinNet with significantly fewer parameters (88.4%) and FLOPs (77.2%).
Conclusions:
- SwinEGNet effectively captures global context and enhances feature fusion for RGB-D SOD.
- The proposed edge-guided approach significantly improves detection accuracy and efficiency.
- The model's performance and efficiency suggest its potential for real-world computer vision applications.

