Related Experiment Video
Updated: Jan 11, 2026

Recording Single Neurons' Action Potentials from Freely Moving Pigeons Across Three Stages of Learning
Published on: June 2, 2014
Bidirectional transition consistency between multi-domain observations for visual reinforcement learning
Xiaobo Hu1, Youfang Lin1, Jinwen Wang1
1Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, School of Computer Science and Technology, Beijing Jiaotong University, Beijing, 100044, China.
Abstract:
Visual reinforcement learning has proven effective in addressing control tasks with high-dimensional image observations. However, obtaining generalizable representations for policy learning under various visual interferences remains a significant challenge. Inspired by human cognitive processes in unfamiliar scenarios, we propose the Multi-Domain Bidirectional Transition (MDBT) model. Unlike prior approaches that either enforce visual consistency or rely solely on precise model-based transitions, MDBT explicitly incorporates multi-domain observations with visual perturbations and kernel regions, while emphasizing task-related dynamics. This design enables MDBT to remove noise interference and preserve task relevance, resulting in more robust and transferable representations.MDBT consists of three key components. Initially, we utilize the Data Transformation module to diversify the original observations, thereby obtaining multi-domain observations with various degrees of visual interference. Subsequently, we employ a Bidirectional Transition module to predict both forward and backward environmental transitions, extracting the task-relevant representations. Finally, we impose a Consistency Target to constrain the coherent multi-domain transition predictions, ensuring that the representations remove noise interference while retaining task relevance. Extensive experiments demonstrate that MDBT achieves consistent state-of-the-art performance: it outperforms prior approaches by an average of 2.3% in success rate on the "Video Hard" setting of the DeepMind Control Suite, improves average generalization by up to 124.1% on the "Reach" task and 73.5% on the "Peg in Box" task of robotic manipulation benchmarks, and enhances average generalization returns by up to 23.2% and average driving distance by up to 16.9% under severe weather and lighting perturbations in CARLA. These results highlight the effectiveness of MDBT in learning robust and transferable representations for visual reinforcement learning.
More Related Videos
Related Concept Videos
Observational Learning
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Associative Learning
Classical conditioning, also known...
Multi-input and Multi-variable systems
In the absence of...
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Reinforcement Schedules
Once a behavior is learned,...

