Related Experiment Videos
DG-SegNet: A depth-guided RGB-D semantic segmentation framework for emergency escape ramps toward traffic accident
Zuosheng Hu1,2, Yun Hao2, Junhao Li3
1School of Electrical Engineering, Southwest Jiaotong University, Chengdu, China.
Objective:
In intelligent transportation scenarios, semantic segmentation plays a crucial role in autonomous driving perception by enabling fine-grained semantic understanding of complex road environments and contributing to the effective reduction of traffic accident risks. However, conventional RGB-based segmentation methods exhibit inherent limitations, and although multimodal information fusion has emerged as a promising direction, existing multimodal models generally suffer from large parameter sizes and high computational complexity.
Methods:
To address the aforementioned challenges, we propose DG-SegNet, an efficient depth-guided RGB-D semantic segmentation framework based on SegFormer. The proposed method introduces depth information exclusively at the shallow C1 feature stage, thereby reducing redundant cross-modal computation, and incorporates a Multi-scale Feature Refinement module together with a Gated Fusion mechanism to structurally reorganize heterogeneous modal features and adaptively regulate cross-modal information flow. The fused multimodal features are subsequently unified to enhance the coherence of the overall feature representation. The dataset was expanded from 411 to 500 samples, with the number of semantic categories increased to nine.
Result:
Our method achieves an MIoU of 78.77% and an MPA of 84.89%, outperforming six classical and state-of-the-art competing approaches evaluated under the same experimental settings. Furthermore, on the public PST900 dataset, comparative experiments against eight advanced methods demonstrate that DG-SegNet attains an MIoU of 84.21% and an MPA of 88.56%, consistently maintaining superior performance across multiple semantic categories.
Conclusions:
DG-SegNet mitigates the semantic representation limitations of single RGB inputs in complex traffic scenarios and provides a feasible solution that balances accuracy and efficiency for effective environmental perception in autonomous driving and intelligent transportation systems, demonstrating strong practical applicability and promising potential for real-world deployment.