Related Experiment Video
Updated: Dec 14, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
887
CASNet: A Cross-Attention Siamese Network for Video Salient Object Detection
Summary
This study introduces CASNet, a novel cross-attention model for video salient object detection. CASNet effectively models spatial-temporal information, outperforming existing methods in accuracy and consistency.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Directly applying image-based models to video salient object detection is challenging due to the need for spatial-temporal information.
- Existing methods struggle with intraframe accuracy and interframe consistency in video saliency detection.
Purpose of the Study:
- To propose a novel cross-attention based encoder-decoder model under the Siamese framework (CASNet) for video salient object detection.
- To improve both intraframe accuracy and interframe consistency in video saliency detection.
Main Methods:
- Developed a Siamese framework incorporating a baseline encoder-decoder backbone trained with Lovász softmax loss.
- Integrated self- and cross-attention modules to preserve saliency correlation and enhance detection consistency.
- Utilized ablation analysis and cross-dataset validation for rigorous evaluation.
Main Results:
- CASNet effectively models spatial-temporal information for video salient object detection.
- The proposed model demonstrates improved intraframe accuracy and interframe consistency.
- Quantitative results show CASNet outperforms 19 state-of-the-art methods on six benchmark datasets.
Conclusions:
- CASNet offers a significant advancement in video salient object detection.
- The cross-attention mechanism is effective in addressing the challenges of spatial-temporal modeling in videos.
- The model's superior performance validates its effectiveness across diverse benchmark datasets.

