对于UAV-PE游戏的等级增强学习,使用替代延迟更新方法.
IEEE transactions on neural networks and learning systems
|February 21, 2024
概括
本研究引入了一种新的等级增强学习 (HRL) 算法,用于无人机追逐逃跑 (UAV-PE) 游戏的替代延迟更新 (ADU) 方法. 该方法通过合动力学和动态学,有效地找到近似的纳什平衡解决方案.
科学领域:
- 机器人和控制系统 机器人和控制系统
- 人工智能的人工智能
- 游戏理论 游戏理论
背景情况:
- 无人驾驶飞行器 (UAV) 追击逃跑 (PE) 游戏提出了复杂的控制挑战.
- 现有的方法经常与这些系统固有的合动力学和动力学相斗争.
- 有效计算纳什平衡 (NE) 解决方案对于开发强大的UAV-PE策略至关重要.
研究的目的:
- 为UAV-PE游戏系统提出一种新的等级增强学习 (HRL) 算法.
- 整合一种替代延迟更新 (ADU) 方法,以提高培训效率.
- 为了获得UAV-PE游戏的近似纳什平衡 (NE) 解决方案,配合动力学和动力学.
主要方法:
- 一个层次化的学习过程,涉及零和动力学和最佳动力学游戏.
- 深度神经网络 (NN) 用于在动态和动态层面近似政策和价值函数.
- 通过确定一个球员的战略来稳定训练的ADU方法.
主要成果:
- 开发一个HRL算法与ADU用于UAV-PE系统.
- 为算法趋同和最佳性推导出足够的条件.
- 获得过载不平等,通过动态控制输入确保动态状态跟踪.
结论:
- 拟议的HRL算法与ADU是可行的和有效的UAV-PE游戏.
- 该方法成功计算了合系统的近似NE解决方案.
- 模拟结果验证了算法的性能和训练效率.
相关概念视频
Reinforcement Schedules
147
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
147
PD Controller: Design
229
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
229
Timing and Consequences on Behavior
94
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
94
Time-Domain Interpretation of PD Control
111
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
111
Feedback control systems
313
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
313


