Related Experiment Video
Updated: Aug 19, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
605
Multi-scale fusion for RGB-D indoor semantic segmentation.
Shiyi Jiang1, Yang Xu2,3, Danyang Li1
1College of Big Data and Information Engineering, Guizhou University, Guiyang, 550025, China.
Scientific Reports
|November 26, 2022
Summary
This study introduces a novel RGB-D indoor semantic segmentation network using wavelet transform to preserve high-frequency details lost in traditional methods. The approach enhances contour accuracy and achieves state-of-the-art real-time performance.
Area of Science:
- Computer Vision
- Image Processing
- Deep Learning
Background:
- Convolution and pooling operations in deep networks often degrade high-frequency information and contour details, particularly in semantic segmentation tasks.
- Existing methods for RGB-D semantic segmentation struggle to fully leverage information from both RGB and depth images.
- Wavelet transform offers a robust method for retaining both low and high-frequency image information.
Purpose of the Study:
- To address the loss of high-frequency information and contour details in RGB-D semantic segmentation.
- To propose an efficient RGB-D indoor semantic segmentation network that effectively utilizes multi-scale and multi-frequency information.
- To improve the accuracy of semantic segmentation, especially for image edge contours.
Main Methods:
- Developed a wavelet transform fusion module to preserve contour details.
- Employed nonsubsampled contourlet transform to replace traditional pooling operations.
- Integrated a multiple pyramid module for aggregating multi-scale and global contextual information.
Main Results:
- The proposed network effectively retains multi-scale information and the complementarity of high and low frequencies using wavelet transform.
- The method improves segmentation accuracy for image edge contours without losing multi-frequency characteristics as network depth increases.
- Achieved state-of-the-art performance and real-time inference on the NYUv2 and SUNRGB-D indoor datasets.
Conclusions:
- The proposed multi-scale fusion network effectively overcomes information loss in RGB-D semantic segmentation.
- Wavelet transform integration enhances the preservation of fine details and contour accuracy.
- The method demonstrates superior performance and efficiency for indoor scene semantic segmentation.

