多种代理自我注意力强化学习,用于多种USV狩猎目标
概括
本研究介绍了一种多代理增强学习方法,使用多头自我注意力为多个无人地面车辆 (多USV) 猎杀目标. 这种新的方法提高了狩猎成功率和算法融合速度.
科学领域:
- 机器人和控制系统 机器人和控制系统
- 人工智能的人工智能
- 海洋工程 海洋工程
背景情况:
- 使用多个无人水面车辆 (多个USV) 的合作目标狩猎带来了重大协调挑战.
- 现有的多代理强化学习 (MARL) 算法可能在复杂场景中难以有效处理信息和快速融合.
研究的目的:
- 开发一种先进的MARL方法,用于多个USV的合作狩猎.
- 通过多USV系统提高目标采集的效率和成功率.
主要方法:
- 建立了USVs的动力学,动态和环境模型.
- 开发了多USV目标狩猎的差异性游戏模型,定义了合作和竞争动态.
- 设计了一个MARL算法,整合了多头自我注意力 (MSA) 机制,以专注于关键信息.
主要成果:
- 拟议的基于MSA的MARL算法显示,与基线MARL相比,合速度提高了17%.
- 在狩猎成功率上实现了8%的增长.
结论:
- 该MSA机制显著提高了MARL算法在多USV合作狩猎任务的性能.
- 开发的方法为自主多USV协调和目标拦截提供了更有效的解决方案.
相关概念视频
Multi-input and Multi-variable systems
93
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
93
Observational Learning
111
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
111
Masking and Demasking Agents
2.3K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.3K
Associative Learning
270
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
270
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Reinforcement
169
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
169


