适应式安全增强学习与全状态约束和自动驾驶汽车受约束的适应
IEEE transactions on cybernetics
|June 26, 2023
概括
本研究介绍了自动驾驶汽车的自适应安全强化学习 (RL) 算法,确保所有状态变量在学习过程中保持在安全范围内. 该方法提高了不确定性下的控制性能和安全性.
科学领域:
- 自主系统 自主系统
- 机器学习 机器学习
- 控制理论 控制理论
背景情况:
- 安全关键的自动驾驶汽车在学习过程中需要状态变量保持在定义区域内.
- 现有的强化学习 (RL) 方法可能无法在整个学习过程中保证安全限制.
研究的目的:
- 为自动驾驶汽车开发一个自适应的安全强化学习 (RL) 算法.
- 确保在整个学习过程中,在安全区域内限制全状态变量.
- 提高自动驾驶汽车系统的控制性能和安全保证.
主要方法:
- 一个自适应的安全RL算法,集成了优化后退技术和不对称的障碍力普诺夫函数 (BLF) 方法.
- 用BLF相关术语和独立学习组件分解子系统控制和价值函数衍生.
- 一个受约束的适应算法与投影操作员来管理安全优化冲突.
主要成果:
- 拟议的算法确保在学习过程中全状态变量保持在安全区域内.
- 与现有方法相比,在自动驾驶汽车运动控制的模拟中证明了卓越的性能.
- 经过验证,收度有所提高,差异减少,尤其是在不确定的条件下.
结论:
- 适应性安全RL算法有效优化了系统控制,同时保证了状态变量约束.
- 该方法通过受约束的适应提供了双重保证的安全性能.
- 该方法被验证用于提高自动驾驶汽车控制系统的安全性和性能.
相关概念视频
Controller Configurations
128
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
128
Avoidance Learning and Learned Helplessness
1.8K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.8K
Constraints and Statical Determinacy
640
In structural engineering, the equilibrium of a system is not only determined by its equations of equilibrium but also with the help of constraints. Constraints refer to restrictions on the motion of a system. The proper combinations of constraints can minimize the total number of constraints needed to maintain a system in mechanical equilibrium. When this happens, the system is said to be statically determinate. For such systems, the unknown reaction supports can be estimated using equilibrium...
640
PD Controller: Design
291
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
291
Rolling Resistance: Problem Solving
385
Rolling resistance, also known as rolling friction, is the force that resists the motion of a rolling object, such as a wheel, tire, or ball, when it moves over a surface. It is caused by the deformation of the object and the surface in contact with each other, as well as other factors like internal friction, hysteresis, and energy losses within the materials. Rolling resistance opposes the object's motion, requiring additional energy to overcome it and maintain movement. In practical...
385
Observational Learning
222
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
222


