基于强化学习的多个四旋翼无人机的形成周围控制多个四旋翼无人机的追逐逃跑游戏
1Logistics Engineering College, Shanghai Maritime University, Shanghai 201306, China.
ISA transactions
|December 17, 2023
概括
本研究介绍了一种强化学习 (RL) 控制方法,用于追逐逃跑游戏中的多个四旋翼无人机 (UAV). 这种方法确保追逐者尽管有外部干扰,但也同样包围了逃避者.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 控制系统 控制系统
- 人工智能的人工智能
背景情况:
- 多重四旋翼无人机 (UAV) 越来越多地用于复杂的场景.
- 追逐-逃避 (PE) 游戏带来了重大的控制挑战,特别是在多个代理和外部干扰的情况下.
- 在多无人机系统中实现稳定的阵列和协调的机动需要先进的控制策略.
研究的目的:
- 开发一种基于强化学习 (RL) 的围绕阵营的控制方法,用于多个无人机的追逐-逃避 (MPE) 游戏.
- 为了应对追逐者同样围绕逃避者的挑战,同时在外部干扰下保持阵营完整性.
- 为了确保四旋翼无人机位置和姿态跟踪错误子系统的稳定性.
主要方法:
- 提出了一种新的控制策略,结合了前控制和强化学习 (RL).
- 开发了两个新的成本功能,适用于面临外部干扰的四旋翼无人机.
- 实施了仅批评神经网络 (NN) 重量更新规律,以实现最佳的成本函数估计.
- 利用RL解决汉密尔顿-雅各比-艾萨克斯 (HJI) 方程以实现纳什平衡.
- 应用欧勒公式来证明周围相等的属性.
主要成果:
- 使用基于RL的控制方案证明了跟踪错误子系统的稳定性.
- 成功实现了多重四旋翼无人机系统的纳什平衡.
- 首次证明了追逐者围绕逃犯者等同周围的属性.
- 数字模拟证实了拟议方法的有效性和卓越性能.
结论:
- 拟议的基于RL的围绕阵列的控制方法对于多个无人机的追逐逃跑游戏是有效的.
- 该方法确保了稳定的控制,并在外部干扰下实现了所需的相等的周围形成.
- 这项研究通过整合RL用于复杂的协调行为来推进多代理控制系统.
相关概念视频
Collisions in Multiple Dimensions: Problem Solving
4.2K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.2K
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Collisions in Multiple Dimensions: Introduction
5.4K
It is far more common for collisions to occur in two dimensions; that is, the initial velocity vectors are neither parallel nor antiparallel to each other. Let's see what complications arise from this. The first idea is that momentum is a vector. Like all vectors, it can be expressed as a sum of perpendicular components (usually, though not always, an x-component and a y-component, and a z-component if necessary). Thus, when the statement of conservation of momentum is written for a...
5.4K
Three-Dimensional Force System:Problem Solving
669
A three-dimensional force system refers to a scenario in which three forces act simultaneously in three different directions. This type of problem is commonly encountered in physics and engineering, where it is necessary to calculate the resultant force on the system, which can then be used to predict or analyze the behavior of the object or structure under consideration.
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
669
Observational Learning
181
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
181
Open and closed-loop control systems
758
Control systems are foundational elements in automation and engineering. They are broadly categorized into open-loop and closed-loop systems. These classifications hinge on the presence or absence of feedback mechanisms, significantly influencing the system's performance, complexity, and application.
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
758


