协作双重参与者框架使用深度决定性政策梯度用于灵活的批量流程
Xindong Wang1, Zidong Liu1, Junghui Chen2
1College of New Energy, China University of Petroleum (East China), Qingdao, 266580, Shandong, China.
概括
本研究引入了一种新的深度强化学习 (DRL) 方法,用于灵活的批处理过程控制. 协作双重行动者方法提高了控制性能,尽管条件不同.
科学领域:
- 过程控制 过程控制
- 人工智能的人工智能
- 化学工程是化学工程的重要组成部分.
背景情况:
- 批量加工是高效的,但具有挑战性,需要灵活的条件来控制.
- 传统的批量对批量学习控制在有限的预先信息中扎.
- 在动态批量系统中优化性能需要先进的控制策略.
研究的目的:
- 开发一种新的深度强化学习 (DRL) 方法,用于灵活的批量过程控制.
- 解决传统方法在处理不同操作条件和初始状态方面的局限性.
- 加强控制政策的制定,确保复杂批次系统的安全运行.
主要方法:
- 提出了一种基于双主角的深度决定性政策梯度 (CTA-DDPG) 的协作方法.
- 利用连续的演员-批评网络与一个共享的批评者,用于离线的元政策探索和在线性能提升.
- 纳入政策整合和时空体验重复,以实现强大的转移和高效的学习.
主要成果:
- CTA-DDPG证明了对灵活的批量流程制定有效的控制政策.
- 该方法确保了在不同的试验长度和初始条件下安全运行.
- 对数值示例和注塑成型工艺的评估证实了卓越的性能.
结论:
- CTA-DDPG方法为灵活的批量流程控制提供了优质的解决方案.
- 这种DRL方法有效地克服了传统学习控制策略的局限性.
- 提出的方法可以在复杂,动态的工业环境中实现所需的控制结果.
更多相关视频
11:54Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
Published on: May 8, 2021
4.3K
12:54Density Gradient Multilayered Polymerization DGMP: A Novel Technique for Creating Multi-compartment, Customizable Scaffolds for Tissue Engineering
Published on: February 12, 2013
12.4K
相关概念视频
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
45
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
45
Parallel Processing
125
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
125
Multi-input and Multi-variable systems
86
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
86
Statically Indeterminate Problem Solving
340
Statically indeterminate problems are those where statics alone can not determine the internal forces or reactions. Consider a structure comprising two cylindrical rods made of steel and brass. These rods are joined at point B and restrained by rigid supports at points A and C. Now, the reactions at points A and C and the deflection at point B are to be determined. This rod structure is classified as statically indeterminate as the structure has more supports than are necessary for maintaining...
340
Associative Learning
239
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
239
