Related Experiment Video
Updated: May 21, 2026

03:31
End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
A visual-attention model using Earth Mover's Distance-based saliency measurement and nonlinear feature combination.
Yuewei Lin1, Yuan Yan Tang, Bin Fang
1Department of Computer Science and Engineering, University of South Carolina, Columbia, SC 29208, USA. ywlin.cq@gmail.com
Summary
This study presents a novel computational visual-attention model using Earth Mover's Distance (EMD) for improved static and dynamic saliency maps. The new model demonstrates superior performance compared to existing methods in visual attention research.
Area of Science:
- Computer Vision
- Computational Neuroscience
- Image Processing
Background:
- Traditional visual attention models often rely on Difference-of-Gaussian filters.
- Combining features in existing models can be suboptimal.
- There is a need for advanced models for both static and dynamic saliency mapping.
Purpose of the Study:
- To introduce a novel computational visual-attention model.
- To improve the accuracy of static and dynamic saliency maps.
- To offer a biologically inspired approach to visual attention.
Main Methods:
- Utilized Earth Mover's Distance (EMD) for center-surround difference calculation, replacing Difference-of-Gaussian filters.
- Implemented a two-step nonlinear feature combination process: Lm-norm for super-feature creation and Winner-Take-All for final combination.
- Extended the model to generate dynamic saliency maps using spatiotemporal receptive fields and EMD.
Main Results:
- The proposed model achieved superior performance on both static image and video datasets.
- EMD proved effective in measuring center-surround differences in visual attention.
- The biologically inspired feature combination enhanced model capabilities.
Conclusions:
- The novel visual-attention model offers a significant advancement over existing methods.
- The use of EMD and specific nonlinear operations provides a more effective approach to saliency mapping.
- The model's success in both static and dynamic scenarios highlights its versatility and robustness.
