数据模型混合驱动的安全增强学习,用于对不安全的移动区域进行自适应避开控制
IEEE transactions on neural networks and learning systems
|April 18, 2025
概括
本研究引入了一种新的安全强化学习 (SRL) 方法,用于避免控制. 该方法可确保复杂环境中的安全,移动不安全区域,提高控制系统的可靠性.
科学领域:
- 机器人和控制系统 机器人和控制系统
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 安全是扩大强化学习 (RL) 应用的关键问题.
- 现有的方法在具有多个移动不安全区域的动态环境中难以避免控制.
研究的目的:
- 开发一种新的数据模型混合动力驱动的安全RL (SRL) 方案,用于有效的避免控制.
- 为了应对在多个,移动不安全区域的领域中运行的挑战.
主要方法:
- 将障碍函数 (BF) 编码到成本函数中,以将避免问题转化为最佳控制问题.
- 使用整体RL (IRL) 具有演员关键神经网络 (NN) 结构和状态跟踪 (StaF) 内核功能,用于自适应性政策生成.
- 采用状态推断技术,将实时和模拟经验整合到政策学习中.
主要成果:
- 关于拟议的SRL方案的闭环稳定性和重量趋同的理论依据.
- 在单个集成器,非线性数值和单轮运动系统上证明了有效性.
- 通过比较分析,突出了对现有控制方法的优势.
结论:
- 拟议的数据模型混合驱动的SRL方案为复杂,动态环境中的避免控制提供了强大的解决方案.
- 该方法确保了安全性和稳定性,同时实现了有效的控制政策学习.
- 这项工作促进了安全RL在现实世界控制场景中的实际应用.
相关概念视频
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
PD Controller: Design
145
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
145
Fixed Action Patterns
15.6K
A fixed action pattern (FAP) is a specific, hard-wired sequence of behaviors that occurs in response to an external stimulus, called a sign stimulus. The behavior is “fixed” because it is essentially unchangeable—proceeding similarly across individuals of a species every time it occurs.
15.6K
Instinctive Drift
153
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
153


