深度强化学习的应用在无人机群组中用于地面监视
Raúl Arranz1, David Carramiñana1, Gonzalo de Miguel1
1Information Processing and Telecommunications Center, Universidad Politécnica de Madrid, ETSI Telecomunicación, Av. Complutense 30, 28040 Madrid, Spain.
Sensors (Basel, Switzerland)
|November 14, 2023
概括
本研究介绍了空中群体的混合人工智能系统,使用深度强化学习来有效地监视区域和跟踪目标. 该系统展示了安全应用程序的高效搜索,目标获取和一致跟踪.
科学领域:
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
- 航空航天工程 航空航天工程
背景情况:
- 目前的空中小群管理依赖于经典和强化学习方法.
- 需要先进的人工智能系统用于监视和执法应用.
研究的目的:
- 提出一种混合人工智能系统,将深度强化学习集成到集中式群体架构中.
- 为了使空中小群能够进行有效的区域监视,地面目标搜索和跟踪.
主要方法:
- 一个混合人工智能系统,具有中央群控制器和深度强化学习 (DRL) 训练的代理人.
- 利用近接政策优化 (PPO) 算法进行代理行为训练.
- 定义的指标来评估群体在监视和跟踪任务中的表现.
主要成果:
- 模拟显示,该系统有效地搜索操作区域.
- 该系统在合理的时间框架内证明了有效的目标获取.
- 实现了实现目标的一致和持续跟踪.
结论:
- 拟议的混合人工智能系统为空中集群式监视提供了有效的解决方案.
- 深度增强学习增强了合作无人机在安全应用中的能力.
- 该系统在搜索,获取和跟踪地面目标方面提供可靠的性能.
相关概念视频
Absolute Motion Analysis- General Plane Motion
Visualize a drone, with its propellers spinning rapidly, hovering mid-air. The fascinating movements and operations of this drone can be comprehended by applying the principle of general plane motion.
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the drone...
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the drone...
Reinforcement
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Vector Functions and Motion: Problem Solving
Accurate position tracking is fundamental to the safe and effective operation of unmanned aerial vehicles (UAVs), particularly during precision maneuvers near complex structures. In this scenario, a drone is programmed to perform a high-precision inspection of a vertical structure, starting at position ((x, y, z) = (3, 0, 0)), with an initial velocity oriented in the positive z-direction. The trajectory of the drone is governed by a time-dependent acceleration function a(t), which is predefined...


