基于Q值和环境认知的强化学习的自适应探索策略
Tenglong Yang1, Jingbao Hou2, Peiyi Zhang3
1National-local Joint Engineering Laboratory of Marine Mineral Resources Exploration Equipment and Safety Technology, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China; College of Mechanical and Electrical Engineering, Hunan University of Science and Technology, Xiangtan, 411201, Hunan, China.
概括
本研究介绍了Var,这是一种适应性探索策略,用于强化学习,平衡探索和开发. 瓦尔可以提高学习效率,减少各种环境中的灾难性行为.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 强化学习是一种强化学习.
背景情况:
- 强化学习 (RL) 在顺序决策方面表现出色,但在勘探和开采平衡方面扎.
- 现有的方法经常使用单个信号,导致效率低下的探索和低于最佳的解决方案.
研究的目的:
- 开发一种适应性探索策略,用于强化学习.
- 通过解决当前勘探技术的局限性来提高学习效率和绩效.
主要方法:
- 提议Var,一种使用Q值和环境认知 (Q值差异和新奖金) 的自适应性勘探策略.
- 将Var集成到Q学习 (Var-QL) 和深度Q网络 (Var-DQN) 中.
- 引入大型表格离散环境的LRU-Var-QL,包括最近使用最少 (LRU) 缓存.
主要成果:
- 瓦尔鼓励最初进行广泛和深入的勘探,然后过渡到开发.
- LRU-Var-QL在FrozenLake-v1.1.上显示了减少的灾难性行为.
- 与基线相比,Var-QL和Var-DQN在Atari游戏中获得了更高的累积回报和更快的学习速度.
结论:
- 拟议的Var战略增强了强化学习的探索.
- 基于Var的方法在不同环境中显示了学习速度和表现的显著改善.
相关概念视频
Cognitive Learning
1.4K
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
1.4K
Decision Making: P-value Method
7.0K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
7.0K
Avoidance Learning and Learned Helplessness
2.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.7K
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Environmental Influences on Intelligence
1.0K
Despite the strong genetic influence on traits like intelligence, environmental factors significantly shape outcomes. For example, while over 90% of height variation is due to genetic differences, environmental factors such as nutrition also have a notable impact. Similarly, for intelligence, changes in a child's surroundings can significantly alter their IQ. Research shows that enriched environments boost children's academic success and help them develop key cognitive skills. Children...
1.0K
Reinforcement
992
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
992


