双头预测和重建使用粗细面具用于视觉增强学习
Yun Zhou1, Yuqiang Wu2, Qiaoyun Wu2
1School of Artificial Intelligence, Anhui University, Hefei, China; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei, China.
概括
采用粗细面具 (DPRM) 的双头预测和重建方法通过改进表示学习来增强视觉增强学习 (RL). 这种方法提高了代理的性能,并加快了复杂的控制任务的趋同.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 机器人技术 机器人技术 机器人技术
背景情况:
- 在具有有限经验的高维环境中,有效的表现学习对于视觉增强学习 (RL) 至关重要.
- 利用代理轨迹数据是改善各种任务中的RL性能的关键.
研究的目的:
- 引入双头预测和重建与粗至细面罩 (DPRM) 方法,以增强视觉RL.
- 在训练期间,提高代理人从其采样轨迹中学习的能力.
主要方法:
- DPRM集成粗细面具与双头预测重建 (DHPR) 架构和基于坐标的空间编码策略 (CSCS).
- CSCS增强了空间信息,以更好地检测运动变化.
- 基于变压器的DHPR使用三重输入令牌 (动作 + 观察状态) 进行双向状态预测和特征重建.
主要成果:
- 在连续 (DeepMind控制套件) 和离散 (Atari) 控制任务中,DPRM显著提高了性能.
- 这种方法导致了更高的奖励积累和更快的收率.
- 实验验证证明了拟议方法的有效性.
结论:
- DPRM方法为视觉RL中的表示学习提供了一个强大的解决方案.
- 它有效地解决了由高维数据和有限经验所带来的挑战.
- 在复杂的控制场景中,DPRM提高了学习效率和任务性能.
相关概念视频
Masking and Demasking Agents
3.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
3.4K
Observational Learning
832
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
832
Associative Learning
1.2K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
1.2K
Reinforcement
830
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
830
