具有杀戮链信息的多目标优化等级增强学习,以提高自主无人群群的弹性
Yingdong Gou1, Siwen Wei2, Kai Xu1
1Northwest Institute of Mechanical and Electrical Engineering, Xianyang, 712099, Shaanxi, China.
概括
本研究介绍了HRL-KCIMOO,这是一个新的层次强化学习框架. 它通过整合知识预训练和动态多目标优化来提高自主无人机群 (AUS) 对敌对攻击的弹性,以提高任务成功.
科学领域:
- 机器人和人工智能 机器人和人工智能
- 多代理系统 多代理系统
- 强化学习是一种强化学习.
背景情况:
- 自主无人机群 (AUS) 需要对抗敌对干扰的弹性,才能成功执行任务.
- 当前的方法在复杂的多代理系统中与一般化和融合作斗争.
- 在AUS中单一目标的强化学习 (RL) 隐藏了权衡,限制了适应性反应.
研究的目的:
- 制定一个强有力的框架,以提高澳大利亚对抗威胁的弹性.
- 通过整合知识和多目标优化来克服传统RL的局限性.
- 提高在动态压力下AUS的适应性反应能力.
主要方法:
- 拟议的HRL-KCIMOO:一个层次化的强化学习框架.
- 使用了预先训练有反差学习,拓结构重建和中心意识排名的图表注意力编码器.
- 采用LSTM实现了高级参与者-关键架构,用于适应性目标权重和分散的低级代理执行.
主要成果:
- 在各种对抗性环境中,HRL-KCIMOO表现优于既定基准.
- 该框架维持了显著更高的任务成功率,即使在严重的群体消耗下.
- 实现了目标的动态调和:快速恢复性,运营连续性和系统稳定性.
结论:
- HRL-KCIMOO有效地提高了AUS的弹性和对抗敌对干扰的任务成功.
- 知识预训练和动态多目标优化的整合对于适应性群体行为至关重要.
- 这种方法为在具有挑战性的环境中强大的自主系统提供了显著的进步.
相关概念视频
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K
Collisions in Multiple Dimensions: Problem Solving
5.2K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
5.2K
Observational Learning
804
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
804
Multi-input and Multi-variable systems
378
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
378
Predator-Prey Interactions
21.0K
Predators consume prey for energy. Predators that acquire prey and prey that avoid predation both increase their chances of survival and reproduction (i.e., fitness). Routine predator-prey interactions elicit mutual adaptations that improve predator offenses, such as claws, teeth, and speed, as well as prey defenses, including crypsis, aposematism, and mimicry. Thus, predator-prey interactions resemble an evolutionary arms race.
21.0K
Reinforcement
804
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
804


