在高维度离散行动空间中最被高估的Q值规范化,用于离线增强学习
IEEE transactions on neural networks and learning systems
|December 19, 2025
概括
对于机器人技术而言,深度增强学习 (DRL) 面临着数据收集和Q值高估的挑战. 我们的新方法,MQR,惩罚过高的Q值,提高稳定性和性能在高维的行动空间.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 深度强化学习 (DRL) 对于机器人操纵至关重要,但受到数据收集成本和风险的阻碍.
- 线下强化学习 (RL) 在现有数据上进行训练,但在高维离散行动空间中遭受Q值高估,影响稳定性.
- 在这些空间中,分布外 (OOD) 行动迅速增加,加剧了Q值的高估.
研究的目的:
- 介绍一个新的离线RL算法,最高估的Q值规范化 (MQR),旨在减轻Q值的高估.
- 提高培训稳定性,防止在机器人操纵的高维离散行动空间中出现政策趋同错误.
- 为了验证MQR在挑战各种环境条件下的机器人任务中的有效性.
主要方法:
- 开发了MQR,一个离线RL算法,专门惩罚了最高估的Q值的动作.
- 实施了有针对性的规范化策略,与现有方法中的统一处罚不同.
- 在模拟和现实环境中对机器人推进和抓取任务进行评估MQR,对象安排各异.
主要成果:
- 在机器人操纵任务中,MQR显著超过了基线算法.
- 在模拟中实现了96.94%的清除率,在真实世界密集配置中达到99.04%.
- 展示了高的行动效率和培训稳定性,突出了MQR的稳定性和可扩展性.
结论:
- MQR有效地减轻了在线RL的高维离散动作空间中的Q值高估值.
- 该算法显示出强大的性能,稳定性和适应性,用于现实世界的机器人操纵.
- 在工业机器人应用中,MQR具有很大的部署潜力.
相关概念视频
Detection of Gross Error: The Q Test
6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K
Decision Making: P-value Method
6.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.8K
State Space Representation
499
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
499
Observational Learning
791
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
791
Time-Domain Interpretation of PD Control
348
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
348
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.1K


