一个数字双驱动灵活的调度方法在一个基于层次的强化学习的人机协作车间
Rong Zhang1, Jianhao Lv1, Jinsong Bao1
1College of Mechanical Engineering, Donghua University, Shanghai, 201620 China.
概括
随着COVID-19的爆发,人们越来越需要灵活的制造. 这项研究提出了人机协作系统和数字双胞胎模型,以提高生产适应性并优化人机任务分配以满足动态需求.
科学领域:
- 制造业 工程 制造工程
- 工业自动化 工业自动化
- 运营研究 运营研究
背景情况:
- 随着COVID-19大流行,医疗用品的需求大幅增加,暴露了当前生产线灵活性和效率的局限性.
- 现有的制造系统难以动态地适应波动的市场需求,特别是在全球卫生危机期间.
研究的目的:
- 开发一种灵活的人机协作生产系统,能够动态地适应市场需求.
- 通过数字双胞胎社区模式和社区间融合,增强制造厂生产灵活性.
- 优化人类和机器在生产过程中的参与,以提高效率和负载平衡.
主要方法:
- 建立了作为"平行社区"的平行生产线,并为智能车间构建了一个数字双胞胎社区模型.
- 开发了一个数字双胞胎驱动的社区内部流程优化算法,利用层次强化学习.
- 优化了人类和机器参与工作的比例,以提高整体生产效率和负载平衡.
主要成果:
- 拟议的人机协作系统显示了提高生产灵活性和适应性.
- 数字双胞胎社区模式促进了融合和互动,提高了制造商店的灵活性.
- 层次增强学习算法有效地优化了动态需求场景的人机任务分配.
结论:
- 由数字双胞胎和层次强化学习驱动的智能调度策略显著提高了生产线适应动态需求和变化的能力.
- 人机协作系统对于在面对全球性颠覆时创建具有弹性和响应能力的制造环境至关重要.
- 数字双胞胎社区模型为增强智能车间灵活性提供了一个可行的框架.
相关概念视频
Reinforcement Schedules
212
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
212
Elaborative Rehearsals
108
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
The effectiveness of...
108
Machines: Problem Solving II
336
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. Consider a lifting tong carrying a 100 kg load. It comprises movable sections DAF and CBG linked together with member AB.
336
Observational Learning
222
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
222
Reinforcement
289
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
289
Purposive Learning
153
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
153


