Adaptive exploration strategy in reinforcement learning based on Q-values and environmental cognition.

Tenglong Yang1, Jingbao Hou2, Peiyi Zhang3

  • 1National-local Joint Engineering Laboratory of Marine Mineral Resources Exploration Equipment and Safety Technology, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China; College of Mechanical and Electrical Engineering, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China.

Summary

This study introduces Var, an adaptive exploration strategy for reinforcement learning that balances exploration and exploitation. Var improves learning performance and reduces catastrophic actions in various environments.

Related Concept Videos

Cognitive Learning01:21

Cognitive Learning

Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.4K
Decision Making: P-value Method01:09

Decision Making: P-value Method

The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
7.0K
Avoidance Learning and Learned Helplessness01:14

Avoidance Learning and Learned Helplessness

Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.7K
Observational Learning01:12

Observational Learning

Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Environmental Influences on Intelligence01:29

Environmental Influences on Intelligence

Despite the strong genetic influence on traits like intelligence, environmental factors significantly shape outcomes. For example, while over 90% of height variation is due to genetic differences, environmental factors such as nutrition also have a notable impact. Similarly, for intelligence, changes in a child's surroundings can significantly alter their IQ. Research shows that enriched environments boost children's academic success and help them develop key cognitive skills. Children...
1.0K
Reinforcement01:23

Reinforcement

Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
992