一个基于深度强化学习的关键参与者框架,用于解决灵活的工作车间安排问题的问题.
1School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201620, China.
Mathematical biosciences and engineering : MBE
|February 2, 2024
概括
本研究介绍了一种新的强化学习 (RL) 方法,用于灵活的工作车间安排问题 (FJSP). RL方法有效地优化了调度,在复杂的制造环境中优于传统方法.
科学领域:
- 运营研究 运营研究
- 人工智能的人工智能
- 制造系统工程 制造系统工程
背景情况:
- 工业4.0推动制造业向定制化和灵活性发展.
- 随着市场需求的不断变化,需要为灵活的工作车间安排问题 (FJSPs) 提供先进的解决方案.
- 传统的调度方法在动态和复杂的制造环境中面临挑战.
研究的目的:
- 提出一种新的强化学习 (RL) 方法来解决灵活的工作车间安排问题 (FJSPs).
- 开发一个关键参与者架构,整合基于价值和基于政策的RL方法,以实现最佳的政策制定.
- 提高制造业的灵活性和效率,以应对工业4.0的需求.
主要方法:
- 使用混合演员-关键强化学习架构.
- 制定了马尔科夫决策过程,包括一个全面的特征集和八个行动集,灵感来自调度规则.
- 设计了奖励功能,以最大限度地减少工作完成时间,并确保约束遵守.
主要成果:
- 拟议的RL框架在标准FJSP基准上与启发式,RL和智能算法相比表现优越.
- 该方法在适应性和效率方面取得了显著的改进,特别是在大规模数据集方面.
- 在模拟中始终优于传统的调度方法.
结论:
- 开发的强化学习方法为复杂的灵活工厂调度问题提供了有效的解决方案.
- 演员-关键架构为动态制造环境提供了最佳的政策.
- 这些发现突显了RL在推进工业4.0制造能力方面的潜力.
相关概念视频
Reinforcement Schedules
148
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
148
Reinforcement
209
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
209
Sequence Networks of Rotating Machines
103
A Y-connected synchronous generator, grounded through a neutral impedance, is designed to produce balanced internal phase voltages with only positive-sequence components. The generator's sequence networks include a source voltage that is exclusively in the positive-sequence network. The sequence components of line-to-ground voltages at the generator terminals illustrate this configuration.
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
103
Observational Learning
175
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
175
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
55
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
55
Statically Indeterminate Problem Solving
378
Statically indeterminate problems are those where statics alone can not determine the internal forces or reactions. Consider a structure comprising two cylindrical rods made of steel and brass. These rods are joined at point B and restrained by rigid supports at points A and C. Now, the reactions at points A and C and the deflection at point B are to be determined. This rod structure is classified as statically indeterminate as the structure has more supports than are necessary for maintaining...
378


