通过Actor-Critic强化学习,通过多个无人水面船的自适应性最佳周围控制
IEEE transactions on neural networks and learning systems
|October 18, 2024
概括
本研究介绍了使用演员关键强化学习 (RL) 对多个无人水面船 (USV) 的最佳周围控制算法. 这种新方法可以将轨迹长度和能源消耗减少20%以上.
科学领域:
- 机器人和控制系统 机器人和控制系统
- 人工智能的人工智能
- 海洋工程 海洋工程
背景情况:
- 协调多艘无人水面船 (USV) 提出了复杂的控制挑战.
- 现有的USV环绕控制方法经常与非线性和优化作斗争.
- 强化学习 (RL) 为动态环境中的自适应控制提供了一种有前途的方法.
研究的目的:
- 为多个USVs开发一个最佳的周围控制算法.
- 利用关键演员的强化学习 (RL) 来优化USV的合并过程.
- 为了确保USV形成控制的稳定性和效率.
主要方法:
- 通过汉密尔顿 - 雅各比 - 贝尔曼 (HJB) 方程制定多个USV最佳周围控制问题.
- 实现一个基于贝尔曼残余错误的网络更新规律的自适应性演员-关键RL控制范式.
- 在二级USV中使用虚拟控制器和实际控制器,以确保最佳的周围控制.
- 使用利亚普诺夫理论函数分析控制器稳定性.
主要成果:
- 拟议的基于RL的演员关键环境控制器成功实现了周围的目标.
- 该算法优化了多个USV的演变过程.
- 与现有控制器相比,观察到轨道长度显著减少9.76%和能源消耗减少20.85%.
结论:
- 开发的演员关键RL算法为多个USV的最佳周围控制提供了有效的解决方案.
- 这种方法通过减少轨迹长度和能源消耗来提高效率.
- 该方法在复杂的USV协调任务中表现出强大的稳定性和最佳性能.
相关概念视频
Behavior Modification
130
Behavioral approaches have often been criticized for ignoring mental processes and focusing solely on observable behavior. However, these approaches provide an optimistic perspective for individuals seeking to change their behaviors. Rather than concentrating on intrinsic personality traits, behavioral approaches suggest that even longstanding habits can be modified by changing the reward contingencies that maintain them.
A real-world application of operant conditioning principles is applied...
A real-world application of operant conditioning principles is applied...
130
Operant Conditioning Intervention
44
Operant conditioning serves as a foundational principle in therapeutic interventions aimed at modifying maladaptive behaviors. Central to this approach is the notion that behaviors, both adaptive and maladaptive, are learned through reinforcement. By analyzing the environmental factors that reinforce problematic behaviors, clinicians can design interventions to weaken these reinforcements and replace maladaptive behaviors with healthier alternatives.
In operant conditioning, behaviors that are...
In operant conditioning, behaviors that are...
44


