在不确定的信息下使用TD3-LSTM强化学习算法对无人机进行智能机动决策
Tongle Zhou1, Ziyi Liu1, Wenxiao Jin1
1College of Automation Engineering, Nanjing University of Aeronautics and Astronautics, Nanjing, China.
Frontiers in robotics and AI
|August 18, 2025
概括
本研究引入了一种新的强化学习方法,使用双延迟深确定性政策梯度 (TD3) 和长短期记忆 (LSTM) 进行智能无人机 (UAV) 在复杂的空中对抗中进行机动决策.
科学领域:
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
- 航空航天工程 航空航天工程
背景情况:
- 无人驾驶飞行器 (UAV) 在空中对抗中面临复杂的决策挑战.
- 现有的方法与这些场景固有的不确定性作斗争.
研究的目的:
- 为无人机在空中对抗中开发智能机动决策方法.
- 解决无人机战斗场景中的复杂性和不确定性.
主要方法:
- 一种双延迟深度决定性政策梯度 (TD3) 长短记忆 (LSTM) 强化学习方法.
- 基于无人机作战能力和3DOF无人机模型的胜利/失败判断模型.
- 一个以模型为导向的状态过渡更新机制,用于持续的行动空间决策.
- 使用瓦瑟斯坦距离和内存名义分布对目标检测噪声的不确定性估计.
主要成果:
- 拟议的TD3-LSTM方法有效地从高维度,不确定的空中对抗情况中提取特征.
- 模拟实验证明了该方法在协助无人机做出机动决策方面的能力.
- 该系统成功地导航复杂的空中战斗场景.
结论:
- 开发的强化学习方法增强了在不确定性下无人机机动作决策.
- 这种方法为智能自主空中战斗提供了强大的解决方案.
- 通过各种模拟实验进行进一步验证,证实了该方法的有效性.
相关概念视频
Absolute Motion Analysis- General Plane Motion
271
Visualize a drone, with its propellers spinning rapidly, hovering mid-air. The fascinating movements and operations of this drone can be comprehended by applying the principle of general plane motion.
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the...
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the...
271
Decision Making
230
Decision-making is a fundamental cognitive process that involves evaluating alternatives and selecting among them. This process can range from simple choices, such as deciding what to wear, to complex decisions, like choosing a major in college or a career path. The complexity of the decision often dictates the approach we use, which can be broadly categorized into two types: automatic and controlled decision-making.
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Automatic decision-making is fast, intuitive, and relies on gut feelings...
230
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Observational Learning
311
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
311
Statically Indeterminate Problem Solving
494
Statically indeterminate problems are those where statics alone can not determine the internal forces or reactions. Consider a structure comprising two cylindrical rods made of steel and brass. These rods are joined at point B and restrained by rigid supports at points A and C. Now, the reactions at points A and C and the deflection at point B are to be determined. This rod structure is classified as statically indeterminate as the structure has more supports than are necessary for maintaining...
494
Multi-input and Multi-variable systems
149
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
149


