基于深度强化学习的多变量合系统控制方法
Jin Xu1, Han Li1, Qingxin Zhang1
1School of Artificial Intelligence, Shenyang Aerospace University, Shenyang 110136, China.
Sensors (Basel, Switzerland)
|November 14, 2023
概括
本研究介绍了复杂的多变量系统的深度强化学习控制方法. 与传统方法相比,这种新方法提高了控制精度和稳定性.
科学领域:
- 控制工程 控制工程 控制工程
- 人工智能的人工智能
- 系统科学 系统科学
背景情况:
- 由于多循环合,传统的控制方法难以精确控制多变量系统.
- 开发先进的控制策略对于提高系统性能和稳定性至关重要.
研究的目的:
- 为多变量合系统提出一种基于深度强化学习的新控制方法.
- 为了实现稳定和准确的控制超出现有技术的能力.
主要方法:
- 利用近接政策优化 (PPO) 算法与tanh激活和正常化的优势函数.
- 重新设计的奖励功能和控制器结构,以适应多变量合系统特性.
- 使用控制量输出的振幅评估控制器性能.
主要成果:
- 拟议的深度强化学习方法与去中心化控制,脱控制和传统的PPO相比,显示出更好的控制效果.
- 在 MATLAB/Simulink 中的模拟验证证实了新控制策略的增强稳定性和精度.
- 重新设计的奖励功能和控制器结构有效地解决了多变量合的挑战.
结论:
- 深度强化学习方法为控制复杂的多变量系统提供了重大进步.
- 这种方法为以前很难用传统技术管理的系统提供了强大而准确的解决方案.
- 该研究强调了针对复杂的工程控制应用程序量身定制的深度强化学习的潜力.
相关概念视频
Multi-input and Multi-variable systems
109
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
109
Open and closed-loop control systems
767
Control systems are foundational elements in automation and engineering. They are broadly categorized into open-loop and closed-loop systems. These classifications hinge on the presence or absence of feedback mechanisms, significantly influencing the system's performance, complexity, and application.
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
767
Feedback control systems
319
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
319
Control System Problem
120
In an open-loop system, such as a basic thermostat, the poles of the transfer function influence the system's response but do not determine its stability. However, when feedback is introduced to form a closed-loop system, such as an advanced thermostat that adjusts heating based on room temperature, stability is governed by the new poles of the closed-loop transfer function.
When forming a closed-loop system, issues can arise if the poles cross into the unstable region, leading to potential...
When forming a closed-loop system, issues can arise if the poles cross into the unstable region, leading to potential...
120
Reinforcement
221
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
221
Reinforcement Schedules
160
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
160


