通过离线强化学习的目标Q值来实现适应性的悲观主义
Jie Liu1, Yinmin Zhang2, Chuming Li2
1The Chinese University of Hong Kong, Shatin, NT, Hong Kong Special Administrative Region of China; Shanghai Artificial Intelligence Laboratory, No. 701, Yunjin Road, Shanghai, China.
概括
离线强化学习 (RL) 方法解决了分布之外的操作中的错误. 通过目标Q值 (APTQ) 新的自适应悲观主义算法通过自适应地平衡约束和目标来改善政策学习.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 机器人技术 机器人技术 机器人技术
背景情况:
- 离线强化学习 (RL) 从固定数据集学习,而无需环境交互.
- 数据集中的分布外 (OOD) 操作可以导致Q值估计中的重大错误.
- 现有的OOD行动处理方法往往缺乏或过度悲观,阻碍了政策学习.
研究的目的:
- 开发一种新的线下RL算法,以适应的方式平衡悲观主义约束与标准RL目标.
- 为了应对不同任务和行为政策中数据分布变化的挑战.
- 提高在线RL环境中政策学习的稳定性和性能.
主要方法:
- 引入了使用Q值量数来表示固定数据集内的值分布的概念.
- 通过目标Q值 (APTQ) 算法设计了自适应悲观主义.
- APTQ通过定位Q值量来适应平衡悲观主义约束和RL目标.
主要成果:
- 拟议的APTQ算法表现出比最先进的方法更好的性能.
- 在D4RL-v0基准指标上实现了6.20%的性能改善.
- 在D4RL-v2基准指标上显示出1.89%的性能改善.
结论:
- Q-value量子值对于理解线下RL数据集中的值分布是有效的.
- APTQ提供了一个稳定和适应性的方法来平衡悲观主义和RL目标.
- 该方法显著提高了在线RL任务中的政策学习性能.
相关概念视频
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


