分散的共识推断为多重约束无人机追逐逃避游戏的基层强化学习
概括
一个新的基于共识推断的层次强化学习 (CI-HRL) 框架改善了多个无人机的合作逃避和形成覆盖. 这种方法提高了复杂的追逐逃跑游戏中的群体协调和任务完成.
科学领域:
- 机器人和控制系统 机器人和控制系统
- 人工智能的人工智能
- 多代理系统 多代理系统
背景情况:
- 多个四旋翼无人机 (UAV) 系统对于追逐逃跑游戏 (MC-PEG) 等应用至关重要.
- 合作逃避和形成覆盖 (CEFC) 是一个具有挑战性的MC-PEG任务,特别是有限的通信.
- 高维的复杂性来自于在障碍物,敌人,目标和阵营动态中协调无人机.
研究的目的:
- 为CEFC在多无人机系统中的任务提出一个新的两级框架,即基于共识推断的层次增强学习 (CI-HRL).
- 在通信有限的环境和高维问题空间中应对挑战.
- 增强无人机群体的协作逃避和任务完成能力.
主要方法:
- 制定了两级框架:CI-HRL,其中一个是针对目标定位的高级政策,另一个是针对导航和形成的低级政策.
- 为高层政策引入了以共识为导向的多代理沟通 (ConsMAC),以实现地方国家的全球认知和共识.
- 采用了基于培训的多个代理的深度决定性政策梯度 (AT-M) 和低级别控制的政策蒸.
主要成果:
- CI-HRL在提高群体的协作逃避和任务完成方面表现出卓越的表现.
- 该ConsMAC模块有效地使代理商能够汇总邻居消息并建立共识.
- 软件在循环 (SITL) 模拟验证了框架在复杂场景中的有效性.
结论:
- 拟议的CI-HRL框架为具有挑战性的多UAV合作逃避和形成覆盖任务提供了强大的解决方案.
- 具有专业沟通和控制政策的等级方法显著改善了群体的表现.
- 这项研究促进了多代理系统在动态和受限制环境中的功能.
相关概念视频
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Observational Learning
317
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
317
Collisions in Multiple Dimensions: Problem Solving
4.4K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.4K
Reinforcement
345
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
345
Three-Dimensional Force System:Problem Solving
863
A three-dimensional force system refers to a scenario in which three forces act simultaneously in three different directions. This type of problem is commonly encountered in physics and engineering, where it is necessary to calculate the resultant force on the system, which can then be used to predict or analyze the behavior of the object or structure under consideration.
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
863
Hydraulic Jump: Problem Solving
146
To analyze a hydraulic jump in a rectangular channel with a flow speed of 6 meters per second, follow these steps:Calculate Effective Upstream Velocity:When the downstream gate closes, a hydraulic jump forms, traveling upstream at 2 meters per second. This wave speed combines with the initial channel flow velocity, creating an effective upstream velocity.Identify Flow Velocities Before and After the Hydraulic Jump:Upstream of the hydraulic jump, the effective flow velocity includes both the...
146


