AHEGC:适应性后视体验重播与目标修改的好奇心模块机器人控制控制
IEEE transactions on neural networks and learning systems
|August 1, 2023
概括
通过目标修改的好奇心模块 (AHEGC) 进行自适应的后视体验重复,在稀疏的奖励环境中改善了强化学习. 这种方法提高了样本和勘探效率,使机器人控制任务的学习速度更快.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 强化学习 (RL) 在机器人控制方面表现有前途,但在奖励函数设计和稀疏奖励环境方面存在困难.
- 追溯经验重复 (HER) 通过重复使用失败的经验来解决稀疏的回报,但通常需要广泛的培训.
- 在复杂的机器人任务中,高效的勘探和样本利用仍然是RL的关键挑战.
研究的目的:
- 引入一种新的技术,即Adaptive HER with Goal-Amended Curiosity (AHEGC),以提高机器人控制RL中的样本和探索效率.
- 通过优化培训数据和探索策略,提高RL代理人在稀疏奖励环境中的学习能力.
- 克服现有的HER方法的局限性,特别是长时间的培训时间.
主要方法:
- 开发了一种适应性策略,以调整事后经验 (HE) 采样率和奖励权重,以提高采样效率.
- 整合了一个好奇心机制,以促进更有效的环境探索.
- 引入了目标修改 (GA) 好奇心模块,以减轻过度寻求新奇的行为.
主要成果:
- 拟议的AHEGC方法在6个具有挑战性的机器人控制任务 (Fetch和Hand环境) 中,与现有方法相比,表现优越.
- 实验证实了学习能力和融合速度的显著改善.
- 适应性抽样和GA好奇心模块在稀疏奖励场景中有效平衡了勘探和开发.
结论:
- 在机器人控制中,AHEGC提供了一种更高效和有效的强化学习方法,特别是在奖励稀少的环境中.
- 该方法成功地提高了样本和勘探效率,从而更快地获得技能.
- 在解决奖励功能设计和机器人RL训练效率方面的挑战方面,AHEGC代表了重大进展.
相关概念视频
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Observational Learning
210
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
210
Avoidance Learning and Learned Helplessness
1.8K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.8K
Purposive Learning
142
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
142
Hydraulic Jump: Problem Solving
83
To analyze a hydraulic jump in a rectangular channel with a flow speed of 6 meters per second, follow these steps:Calculate Effective Upstream Velocity:When the downstream gate closes, a hydraulic jump forms, traveling upstream at 2 meters per second. This wave speed combines with the initial channel flow velocity, creating an effective upstream velocity.Identify Flow Velocities Before and After the Hydraulic Jump:Upstream of the hydraulic jump, the effective flow velocity includes both the...
83


