Related Experiment Video
Updated: May 22, 2025

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
Rethinking exploration-exploitation trade-off in reinforcement learning via cognitive consistency
1Key Laboratory of Computational Intelligence and Chinese Information Processing of Ministry of Education and the School of Computer and Information Technology, Shanxi University, Taiyuan 030006, Shanxi, China.
This study introduces a Cognitive Consistency (CoCo) framework for deep reinforcement learning (RL). CoCo enhances sample efficiency and performance by maintaining cognitive consistency through pessimistic exploration.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Deep Reinforcement Learning
Background:
- The exploration-exploitation dilemma is a core challenge in deep reinforcement learning (RL).
- Existing exploration methods often lack sample efficiency.
- Humans maintain cognitive consistency for effective decision-making.
Purpose of the Study:
- To propose a novel framework, Cognitive Consistency (CoCo), for deep reinforcement learning.
- To improve sample efficiency and performance in RL by rethinking the exploration-exploitation trade-off.
- To leverage cognitive consistency principles for more intelligent agent behavior.
Main Methods:
- Developed the Cognitive Consistency (CoCo) framework.
- Utilized a self-imitating distribution correction for optimal policy cognition.
- Implemented pessimistic exploration with inconsistency-minimization objectives inspired by label distribution learning.
Main Results:
- The CoCo framework was validated on standard off-policy RL tasks.
- Maintaining cognitive consistency demonstrably improved sample efficiency.
- Enhanced performance was observed across tested RL tasks.
Conclusions:
- Cognitive consistency offers a promising approach to enhance RL.
- The CoCo framework effectively balances exploration and exploitation.
- This method improves both the speed and quality of learning in RL agents.
Related Concept Videos
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Cognitive Dissonance
The Anchoring-and-Adjustment Heuristic
Instinctive Drift
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...

