政策之外的二维强化学习,以优化批量流程的跟踪控制,以网络诱导的脱落和干扰
Xueying Jiang1, Min Huang1, Huiyuan Shi2
1College of Information Science and Engineering, Northeastern University, China; State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, China.
ISA transactions
|November 29, 2023
概括
本研究介绍了一种非政策的二维强化学习方法,用于批处理过程中的最佳跟踪控制,解决网络问题,如网络中断. 新方法提高了控制精度,尽管存在系统不确定性和干扰.
科学领域:
- 控制工程 控制工程 控制工程
- 人工智能的人工智能
- 化学工程是化学工程的重要组成部分.
背景情况:
- 批量流程带来复杂的控制挑战,网络引起的问题如数据丢失和外部干扰加剧了这一问题.
- 优化跟踪控制 (OTC) 对于在这些动态系统中保持所需的性能至关重要.
- 现有的控制方法可能会与二维 (2D) 批处理和网络不确定性的独特特征作斗争.
研究的目的:
- 为批处理过程中的OTC提出一种新的政策之外的2D强化学习方法.
- 制定一个强有力的控制策略,有效地处理网络引起的停机和干扰.
- 为了提高系统状态的估计,并优化2D批次系统的控制性能.
主要方法:
- 一个放弃的2D增强史密斯预测器被设计用于使用历史时间和批量数据进行状态估计.
- 该研究定义和分析了掉落2D值和Q函数,推导出掉落2D贝尔曼方程.
- 介绍了两个算法:一个离线的2D政策代和一个离线的2D Q学习算法.
主要成果:
- 开发了非政策的2D Q学习算法,仅使用系统输入和估计状态.
- 分析证实了解决方案的公正性和拟议算法的融合性质.
- 通过模拟批量填充过程来证明方法的有效性.
结论:
- 拟议的非政策的二维强化学习方法为OTC在批处理过程中提供了可行的解决方案,其中包括失效和干扰.
- 开发的算法通过有效估计状态和学习最佳政策来提供强大而准确的控制.
- 该研究验证了新型控制策略在工业批量运营中的实际适用性.
相关概念视频
Time-Domain Interpretation of PD Control
119
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
Consider the example of control of motor torque. Initially, a positive...
119
PD Controller: Design
241
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
241
Feedback control systems
316
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
316
Control Systems
1.2K
Control systems are everywhere in contemporary society, influencing diverse applications from aerospace to automated manufacturing. These systems can be found naturally within biological processes, such as blood sugar regulation and heart rate adjustment in response to stress, as well as in man-made systems like elevators and automated vehicles. A control system is essentially a network of subsystems and processes that collaboratively convert specific inputs into desired outputs.
At the heart...
At the heart...
1.2K
Reinforcement Schedules
149
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
149
Second Order systems II
113
In an underdamped second-order system, where the damping ratio ζ is between 0 and 1, a unit-step input results in a transfer function that, when transformed using the inverse Laplace method, reveals the output response. The output exhibits a damped sinusoidal oscillation, and the difference between the input and output is termed the error signal. This error signal also demonstrates damped oscillatory behavior. Eventually, as the system reaches a steady state, the error diminishes to zero.
113


