走向基于安全增强学习的自动公路驾驶的强有力的决策
Rui Zhao1, Ziguo Chen1, Yuze Fan1
1College of Automotive Engineering, Jilin University, Changchun 130025, China.
Sensors (Basel, Switzerland)
|July 13, 2024
概括
本研究引入了安全自动驾驶的新框架,使用重复缓冲器受限制政策优化 (RECPO). 该方法增强了强化学习政策,以实现高速公路驾驶的稳定性,实现零碰撞.
科学领域:
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
- 自主系统 自主系统
背景情况:
- 强化学习 (RL) 对自动驾驶有效,但在各种数据中,特别是长尾场景中,它与强大的安全性斗争.
- 确保安全性和考虑数据分布变化是自动驾驶汽车基于RL的决策的关键挑战.
研究的目的:
- 提出高速公路自动驾驶的新框架,优先考虑安全性和稳健性.
- 开发一种更新RL策略的方法,在遵守安全约束的同时最大限度地提高奖励.
主要方法:
- 引入了重复缓冲器受约束政策优化 (RECPO),以在安全约束范围内更新RL策略.
- 使用重要性抽样和重复缓冲器用于数据重复使用,减轻灾难性遗忘.
- 制定了这个问题作为一个受约束的马尔科夫决策过程 (CMDP) 以优化政策.
主要成果:
- 在CARLA模拟中,RECPO框架显著提高了模型的融合速度,安全性和决策稳定性.
- 在高速公路自动驾驶场景中实现零碰撞率.
- 超过了传统的CPO,深度决定性政策梯度 (DDPG) 和智能驾驶模型+MOBIL (IDM+MOBIL) 方法.
结论:
- 拟议的RECPO方法为培训自动驾驶政策提供了一个强大而安全的方法.
- 该框架有效地解决了自动驾驶的RL安全挑战,特别是在复杂的高速公路环境中.
相关概念视频
Decision Making: Traditional Method
4.0K
The process of hypothesis testing based on the traditional method includes calculating the critical value, testing the value of the test statistic using the sample data, and interpreting these values.
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
4.0K
Reinforcement Schedules
140
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
140
PD Controller: Design
215
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
215
Decision Making: P-value Method
5.3K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.3K
Decision Making
106
Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Automatic decision-making is fast, intuitive, and relies on gut feelings...
106
The Anchoring-and-Adjustment Heuristic
7.2K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.2K


