Related Experiment Video
Updated: Feb 24, 2026

Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions
Published on: May 2, 2019
Adaptive exploration strategy in reinforcement learning based on Q-values and environmental cognition.
Tenglong Yang1, Jingbao Hou2, Peiyi Zhang3
1National-local Joint Engineering Laboratory of Marine Mineral Resources Exploration Equipment and Safety Technology, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China; College of Mechanical and Electrical Engineering, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China.
This study introduces Var, an adaptive exploration strategy for reinforcement learning that balances exploration and exploitation. Var improves learning performance and reduces catastrophic actions in various environments.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Reinforcement Learning
Background:
- Reinforcement learning (RL) excels in sequential decision-making but struggles with exploration-exploitation balance.
- Existing methods often use single signals, leading to inefficient exploration and suboptimal solutions.
Purpose of the Study:
- To develop an adaptive exploration strategy for reinforcement learning.
- To improve learning efficiency and performance by addressing limitations of current exploration techniques.
Main Methods:
- Propose Var, an adaptive exploration strategy using Q-values and environmental cognition (Q-value difference and novelty bonus).
- Integrate Var into Q-learning (Var-QL) and Deep Q-Networks (Var-DQN).
- Introduce LRU-Var-QL for large tabular discrete environments, incorporating a Least Recently Used (LRU) cache.
Main Results:
- Var encourages broad and deep exploration initially, transitioning to exploitation.
- LRU-Var-QL demonstrated reduced catastrophic actions on FrozenLake-v1.
- Var-QL and Var-DQN achieved higher cumulative returns and faster learning on Atari games compared to baselines.
Conclusions:
- The proposed Var strategy enhances reinforcement learning exploration.
- Var-based methods show significant improvements in learning speed and performance across diverse environments.
Related Concept Videos
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Observational Learning
Environmental Influences on Intelligence
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:

