Related Experiment Video
Updated: Sep 9, 2025

Visual Classical Conditioning in Wood Ants
Published on: October 5, 2018
Visual reinforcement learning via sequential consistency preserved policy contrast from optimal transport view
Zehua Zang1, Jiangmeng Li2, Chuxiong Sun2
1University of Chinese Academy of Sciences, Beijing, 101408, China; Institute of Software Chinese Academy of Sciences, Beijing, 100190, China.
Abstract:
Data inefficiency has long posed a significant challenge in the application of visual reinforcement learning methods to complex scenarios. To address this issue, recent studies have incorporated representation learning mechanisms to extract discriminative features by introducing auxiliary objectives that contrast pixel observations. However, our investigations suggest that these representations may not sufficiently capture the essential information for effective decision-making and could potentially impede policy learning. To tackle these limitations, we propose a novel methodology termed CoCo (sequential Consistency preserved policy Contrast). Unlike existing approaches, CoCo emphasizes the capture of invariant policy-based discriminative features by performing policy contrast across multiple distorted views of observations. Accordingly, we determine that there exists a certain intrinsic heterogeneity between policy and observation, since policy establishes an explicit distribution characteristic rather than a plain tensor. To this end, we propose to model the policy contrast as an optimal transport problem and further perform the alignment of policy distributions during contrastive learning. Subsequently, we introduce an inverse consistency-weighting mechanism designed to accentuate the differences between views while maintaining semantic integrity. We establish the theoretical optimality of our proposed method through an information-theoretic analysis and demonstrate its practical effectiveness via comprehensive evaluation across diverse data efficiency benchmarks, where CoCo consistently outperforms existing approaches.
More Related Videos
Related Concept Videos
Observational Learning
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Associative Learning
Classical conditioning, also known...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Purposive Learning

