Related Experiment Video
Updated: May 24, 2026

03:31
End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
CMFGP-Net: RGBT salient object detection based on cross-modal feature global perception and multi-scale deformable
Lijuan Shi1,2, Haitang Li3, Qiuju Liu1,2
1College of Information Engineering, Zhengzhou University of Technology, Zhengzhou, 450052, China.
Scientific Reports
|May 22, 2026
Summary
This study introduces CMFGP-Net, a novel framework for RGB-T salient object detection. It enhances multi-modal feature interaction and robustness against deformation and occlusion, outperforming current methods.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- RGB-T salient object detection faces challenges like poor multi-modal interaction, spatial-spectral misalignment, and lack of robustness to target deformation and occlusion.
- Existing methods struggle to effectively integrate information from visible light and thermal infrared sensors.
Purpose of the Study:
- To propose a robust and effective RGB-T salient object detection framework, CMFGP-Net.
- To address the limitations of insufficient multi-modal feature interaction, spatial-spectral misalignment, and poor robustness in current models.
Main Methods:
- The proposed CMFGP-Net utilizes a Swin Transformer backbone for enhanced global context modeling.
- A cross-modal feature global perception (CMFGP) module is introduced to mitigate spatial-spectral misalignment using temporal correlation and dynamic feature calibration.
- A multi-scale deformable convolutional cross-fusion (MSDCIF) module enhances adaptability to target deformation, occlusion, and complex backgrounds via deformable convolutions.
Main Results:
- CMFGP-Net demonstrates superior performance on VT821, VT1000, and VT5000 datasets, exceeding state-of-the-art methods in E-measure, weighted F-measure, and MAE.
- Ablation studies confirm the significant contribution of each proposed module to the overall performance improvement.
- The model exhibits enhanced robustness and generalization capabilities, particularly in challenging scenarios with complex backgrounds and occlusions.
Conclusions:
- CMFGP-Net effectively addresses key challenges in RGB-T salient object detection.
- The proposed framework offers improved accuracy, robustness, and generalization for salient object detection using multi-modal data.