Related Experiment Video
Updated: Sep 6, 2026

Automated Visual Cognitive Tasks for Recording Neural Activity Using a Floor Projection Maze
Published on: February 20, 2014
Similarity-guided state attention for visual reinforcement learning
Yinfeng Zeng1, Xuesong Wang1, Yuhu Cheng1
1School of Information and Control Engineering, China University of Mining and Technology, Xuzhou, Jiangsu, 221116, China.
Abstract:
Visual reinforcement learning (VRL) aims to extract effective visual information from high-dimensional observations to optimize decision-making policies. While existing VRL methods have achieved significant progress in various control tasks through data augmentation and auxiliary tasks, agents remain susceptible to distractions from redundant information and irrelevant factors, leading to overfitting and degraded generalization in unseen environments. To address this challenge, we propose a similarity-guided state attention (SSA) for VRL. In this method, a similarity guidance module (SGM) is designed to leverage the similarity between state embeddings from the original and augmented observations to guide the encoder to focus on task-relevant regions in the original observation, thereby producing a corresponding state attention map. Meanwhile, the state attention map of augmented observation is obtained by decoding its state embedding. Furthermore, cosine similarity is introduced to measure the global similarity between state attention maps of original and augmented observations, which is incorporated into self-supervised learning objective together with binary cross-entropy loss to encourage alignment of state attention maps in representation space. The combination of SGM and cosine similarity-based alignment of state attention maps facilitates self-supervised learning to obtain more robust state representations for downstream reinforcement learning, thereby enabling the agent to learn optimal policies and improve its generalization ability in unseen environments. Experimental results on the DeepMind Control Generalization Benchmark (DMControl-GB) demonstrate that SSA achieves superior robustness and generalization performance compared with representative VRL baselines.
