Related Experiment Video
Updated: Aug 6, 2026

End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
S2ANet: Semantic-spatial driven alignment salient object detection network in UAV-based unregistered RGB-T image
Abstract:
Detecting salient objects in complex traffic environments remains a challenging task for ground-based systems. Recently, unmanned aerial vehicles (UAVs) have emerged as an effective solution by offering flexible perspectives and the ability to capture complementary RGB and thermal images. However, direct fusion of these modalities often introduces artifacts caused by spatial misalignment and semantic discrepancies. To overcome these challenges, we propose a Semantic-Spatial Driven Alignment Network (S2ANet) for salient object detection in UAV-based unregistered RGB-T imagery. The proposed network performs progressive semantic and spatial alignment through a Semantic-Spatial Alignment (SSA) module, achieving precise cross-modal registration. After alignment, three specialized components are designed to enhance feature representation: the Deep Positional Awareness (DPA) module extracts accurate positional cues via self-attention; the Cross-Hierarchical Contextual Interaction (CCI) module captures both intra- and inter-feature dependencies; and the Multi-scale Detail Perception (MDP) module refines fine-grained details through multi-receptive-field convolutions and spatial attention. Finally, a Dual-Cascade Feature Fusion (DFF) module integrates positional, contextual, and detailed information to generate high-quality saliency maps. Extensive experiments validate that S2ANet achieves accurate feature alignment and superior detection performance in UAV-based RGB-T scenarios.