改进了双DQN与深度强化学习,用于UAV室内自主避障障碍
Ruiqi Yu1, Qingdang Li2, Jiewei Ji1
1College of Data Science, Qingdao University of Science and Technology, Qingdao, 266061, China.
Scientific reports
|August 1, 2025
概括
本研究介绍了一种改进的双深Q网络 (DQN) 算法,用于在无人机中增强自主避障. 这种新的方法在复杂的室内环境中显著提高了安全飞行性能.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 计算机科学 计算机科学
背景情况:
- 无人驾驶飞行器 (UAV) 在复杂的室内环境中面临自主避难障碍的挑战.
- 现有的深度强化学习算法可能缺乏足够的感知和学习能力来应对这些场景.
研究的目的:
- 提出改进的双深Q网络 (DQN) 算法,以提高无人机自主避障性能.
- 优化网络模型并实施动态勘探策略,以提高效率和融合.
主要方法:
- 开发了一种改进的双DQN算法,集成了优化的网络架构和动态探索策略.
- 利用AirSim和虚幻引擎4 (UE4) 创建多种室内模拟环境进行测试.
- 在不同复杂度的场景中评估绩效.
主要成果:
- 在更简单的场景中,平均累积奖励增加了22.88%,平均安全飞行距离增加了23.17%.
- 在复杂的场景中,平均累积奖励增加了2.66%,平均安全飞行距离增加了2.05%.
- 在这两种情景中,最大奖励和安全飞行距离都得到了显著的改善.
结论:
- 提议的改进的双DQN算法有效地提高无人机在复杂的室内环境中自主避障.
- 优化网络和动态勘探战略有助于提高性能和效率.
- 这些发现表明,该算法在需要强大的室内导航的现实应用中具有潜力.
相关概念视频
Absolute Motion Analysis- General Plane Motion
273
Visualize a drone, with its propellers spinning rapidly, hovering mid-air. The fascinating movements and operations of this drone can be comprehended by applying the principle of general plane motion.
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the...
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the...
273
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Collisions in Multiple Dimensions: Problem Solving
4.4K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.4K
One-Degree-of-Freedom System
556
In mechanical engineering, one-degree-of-freedom systems form the basis of a wide range of electrical and mechanical components. Using these models, engineers can predict the behavior of various parts in a larger system, which gives them insight into how different forces interact with each other.
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
556
Controller Configurations
149
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
149
Hydraulic Jump: Problem Solving
143
To analyze a hydraulic jump in a rectangular channel with a flow speed of 6 meters per second, follow these steps:Calculate Effective Upstream Velocity:When the downstream gate closes, a hydraulic jump forms, traveling upstream at 2 meters per second. This wave speed combines with the initial channel flow velocity, creating an effective upstream velocity.Identify Flow Velocities Before and After the Hydraulic Jump:Upstream of the hydraulic jump, the effective flow velocity includes both the...
143


