Related Experiment Video
Updated: Feb 24, 2026

Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions
Published on: May 2, 2019
Adaptive exploration strategy in reinforcement learning based on Q-values and environmental cognition
Tenglong Yang1, Jingbao Hou2, Peiyi Zhang3
1National-local Joint Engineering Laboratory of Marine Mineral Resources Exploration Equipment and Safety Technology, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China; College of Mechanical and Electrical Engineering, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China.
Abstract:
Reinforcement learning has achieved strong results across diverse sequential decision making problems, but balancing exploration and exploitation to avoid local optima remains a core challenge. Many methods rely on a single signal, such as visitation counts, which yields indiscriminate exploration and fails to focus on key decisions with large value differences. We propose Var, an adaptive exploration strategy composed of two parts: Q-values and environmental cognition. Environmental cognition consists of two signals, a Q-value difference term and a novelty bonus. Using these two signals, Var encourages broad and deep exploration early in training and then transitions to an exploitation phase driven by Q-values. To apply the strategy across environments with different state-space characteristics, we integrate it into classical algorithms and introduce Var-QL and Var-DQN. For large tabular discrete environments, we propose LRU-Var-QL, which augments Var-QL with a Least Recently Used (LRU) cache to avoid catastrophic actions. On FrozenLake-v1, LRU-Var-QL reduces catastrophic actions; across nine Atari games, the proposed methods achieve higher cumulative returns and faster learning than strong exploration baselines.
Related Concept Videos
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Observational Learning
Environmental Influences on Intelligence
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:

