基于MachineRank算法和强化学习的灵活工作车间的动态调度
1College of Mechanical and Energy Engineering, Beijing University of Technology, Beijing, 100124, China.
Scientific reports
|November 29, 2024
概括
本研究介绍了一种对决双深Q网络 (D3QN),以优化动态灵活工作室调度问题 (DFJSP),考虑现实世界的干扰. D3QN有效地减少了完成时间,并提高了按时交货率.
科学领域:
- 运营研究 运营研究
- 人工智能的人工智能
- 制造系统工程 制造系统工程
背景情况:
- 动态灵活工厂调度问题 (DFJSP) 由于工作插入,机器故障和处理时间变化等不可预测事件,提出了重大挑战.
- 现有的调度方法往往难以适应现代制造环境的持续和动态性质,尤其是在结合自动引导车辆 (AGV) 时.
研究的目的:
- 开发一个智能调度系统,能够动态适应DFJSP中断.
- 尽量减少最大完成时间 (makespan),并在动态的工作场所环境中提高按时完成率.
主要方法:
- 使用双重深度Q网络 (D3QN) 来学习从连续生产状态中最佳的调度规则.
- 为了提高解决方案质量,提出了一个MachineRank (MR) 算法,导致七个复合调度规则.
- 定义了八个一般状态特征来表示D3QN输入的调度状态.
主要成果:
- 在众多测试实例中,D3QN与各种复合规则,先进的调度规则和标准的Q学习代理相比表现优越.
- 拟议的动态调度触发规则在实时决策中的有效性和合理性得到了验证.
- D3QN方法在不同的生产配置中显示出强烈的普遍性.
结论:
- 基于D3QN的方法为DFJSP提供了有效和强大的解决方案,优于传统方法.
- 连续状态特征和高级强化学习技术的集成为动态调度优化提供了一个有希望的方向.
- 该研究验证了开发的模型在复杂,破坏性制造环境中的实际适用性.
相关概念视频
Reinforcement Schedules
132
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
132
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
41
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
41
Role of Shaping in Operant Conditioning
267
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
267
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K


