通过使用非脆弱的深度强化学习来实现交换飞行器的异步有限时间强大的跟踪控制
Haoyu Cheng1, Ruijia Song2, Haoran Li1
1Unmanned System Research Institute, Northwestern Polytechnical University, Xi'an, China.
Frontiers in neuroscience
|January 8, 2024
概括
一种新的非脆弱的深度强化学习 (DRL) 方法提高了切换无人飞行器的控制能力. 这种方法通过将DRL与强大的控制技术相结合来提高准确性和稳定性,以实现有限时间稳定性.
科学领域:
- 航空航天工程 航空航天工程
- 控制系统 控制系统
- 人工智能的人工智能
背景情况:
- 交换式无人飞行器需要先进的控制以获得准确性和稳定性.
- 传统的控制方法面临着模型不确定性和暂时性能的挑战.
- 深度增强学习 (DRL) 提供了智能控制的潜力,但需要仔细设计稳定性.
研究的目的:
- 提出一种新的非脆弱的深度强化学习 (DRL) 方法,用于切换无人飞行器的有限时间控制.
- 通过将传统的强有力的控制与DRL集成来提高控制准确性,稳定性和智能性.
- 用先进的分析技术在异步切换下确保有限时间稳定.
主要方法:
- 一个混合跟踪控制器,结合了基于动态的控制器 (使用线性矩阵不等式) 和基于学习的控制器 (使用DRL).
- 使用多个Lyapunov函数和取决于模式的平均停留时间的有限时间稳定性分析.
- 在线优化是以马尔科夫决策过程的形式制定的,通过自适应的深度决定性政策梯度算法来解决.
- 在DRL框架内纳入非脆弱控制理论和自适应奖励函数.
主要成果:
- 拟议的DRL方法有效地弥补了模型的不确定性,并提高了过渡控制的准确性.
- 对比模拟表明,所介绍的算法比传统方法的效率更高.
- 综合分析技术证实了飞行器具有异步切换的有限时间稳定性.
结论:
- 新型非脆弱的DRL方法为切换无人驾驶飞行器提供了增强的控制性能.
- 强有力的控制和DRL的整合提供了一个强大的策略来提高系统的准确性和稳定性.
- 开发的算法实现了卓越的稳定性和训练效率,为更智能的自主系统铺平了道路.
相关概念视频
Open and closed-loop control systems
753
Control systems are foundational elements in automation and engineering. They are broadly categorized into open-loop and closed-loop systems. These classifications hinge on the presence or absence of feedback mechanisms, significantly influencing the system's performance, complexity, and application.
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
753
Feedback control systems
314
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
314
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Observational Learning
179
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
179
One-Degree-of-Freedom System
490
In mechanical engineering, one-degree-of-freedom systems form the basis of a wide range of electrical and mechanical components. Using these models, engineers can predict the behavior of various parts in a larger system, which gives them insight into how different forces interact with each other.
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
490
Phase-lead and Phase-lag Controllers
171
Understanding the working function of different types of controllers can be illustrated with practical analogies, such as adjusting a stereo's volume equalizer. Cranking up the bass involves a phase-lead controller, which functions as a high-pass filter, while increasing the treble uses a phase-lag controller, which acts as a low-pass filter. PD controllers, similar to high-pass filters, enhance the system's response to high-frequency components. PI controllers, akin to low-pass...
171


