相关实验视频
Updated: Jul 12, 2025

10:56
Long-term Behavioral Tracking of Freely Swimming Weakly Electric Fish
Published on: March 6, 2014
12.6K
在水下目标狩猎任务中以差异游戏为基础的深度强化学习
IEEE transactions on neural networks and learning systems
|October 27, 2023
概括
本研究介绍了一种新的控制策略,用于多个无人水下车辆 (UUV) 在具有挑战性的海洋条件下猎杀敏捷目标. 拟议的方法提高了在动态,敌对环境中的成功率和稳定性.
科学领域:
- 机器人和控制系统 机器人和控制系统
- 海洋工程 海洋工程
- 游戏理论 游戏理论
背景情况:
- 水下目标狩猎是复杂的,因为动态的海洋环境和对抗目标.
- 现有的方法往往忽略了环境因素,如电流,风和通信延迟.
- 对多个无人水下车辆 (UUV) 的协调控制对于有效的狩猎任务至关重要.
研究的目的:
- 开发UUV的自适应控制方案,以确保一致的目标追逐,并防止在没有碰撞的情况下逃跑.
- 为了利用差异性游戏理论来分析猎人-目标的互动.
- 为了应对水下作业中环境动态和通信限制所带来的挑战.
主要方法:
- 利用差分游戏理论来模拟UUV和高机动性目标之间的对抗行为.
- 用莱布尼茨的公式构想了一个哈密尔顿函数来导出反控制政策.
- 设计了一种修改的多代理强化学习 (MARL) 算法,其中包含了能量流和声学传播延迟.
主要成果:
- 拟议的控制政策确保目标狩猎系统在平均水平上是异常稳定的.
- 该系统实现了纳什平衡,为所有代理商保证了最佳策略.
- 经过修改的MARL方案在模拟中显示出与典型的MARL算法相比更高的性能,实现了更高的奖励和成功率.
结论:
- 开发的控制策略有效地解决了在动态和敌对环境中潜水目标狩猎的复杂性.
- 游戏理论和自适应性控制的整合为多个UUV协调提供了一个强大的框架.
- 修改后的MARL方法为实时轨迹调度和在具有挑战性的水下场景中分布式协调提供了有希望的解决方案.
相关概念视频
Observational Learning
188
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
188
Buoyancy and Stability for Submerged and Floating Bodies
1.8K
In fluid mechanics, buoyancy and stability are key concepts for understanding the behavior of submerged and floating bodies. When a stationary body is fully or partially submerged in a fluid, the fluid exerts a force on the body known as the buoyant force. This force acts vertically upward through a point called the center of buoyancy, which is the center of the displaced fluid volume. According to Archimedes' principle, the magnitude of the buoyant force is equal to the weight of the fluid...
1.8K
Reinforcement
221
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
221

