基于强化学习的预定义时间跟踪控制在标识符-关键-行为体结构下的非线性系统
IEEE transactions on cybernetics
|August 2, 2024
概括
本研究引入了一种新的强化学习控制方法,用于面临干扰的非线性系统. 它实现了快速跟踪错误的融合,可调节的融合时间,确保了系统的稳定性.
科学领域:
- 控制理论 控制理论
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
背景情况:
- 非线性系统容易受到外部干扰,使精确的控制变得复杂.
- 现有的控制方法可能缺乏对收时间或适应未知的动态的能力的保证.
研究的目的:
- 为非线性系统开发一种基于强化学习的新型控制方案.
- 为了实现预定义的时间跟踪控制与规定的性能,尽管外部干扰.
- 为了提高控制灵活性,允许可调节的趋同时间限制.
主要方法:
- 在适应式控制器设计的标识符-关键-行为体框架内使用后退策略.
- 使用神经网络来学习未知的非线性系统动态和控制行为.
- 将规定的性能控制与预定义时间控制概念集成.
主要成果:
- 拟议的方案有效地将追踪错误限制在预定义的附近.
- 收时间的上限可以通过控制参数独立调整.
- 在预定义的时间内,所有系统状态的局限性是使用稳定性理论证明的.
结论:
- 新的强化学习控制方案为非线性系统提供了更好的性能.
- 该方法展示了快速收和可调节的时间限制,优于以前的方法.
- 通过数值和单环操纵器示例验证,展示实际应用.
相关概念视频
Feedback control systems
300
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
300
Linear time-invariant Systems
238
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
238
Control Systems
1.1K
Control systems are everywhere in contemporary society, influencing diverse applications from aerospace to automated manufacturing. These systems can be found naturally within biological processes, such as blood sugar regulation and heart rate adjustment in response to stress, as well as in man-made systems like elevators and automated vehicles. A control system is essentially a network of subsystems and processes that collaboratively convert specific inputs into desired outputs.
At the heart...
At the heart...
1.1K
Reinforcement Schedules
139
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
139
Time-Domain Interpretation of PD Control
87
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
87
Linear Approximation in Time Domain
76
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
76


