Related Experiment Video
Updated: Jul 16, 2025

08:45
A Dual Task Procedure Combined with Rapid Serial Visual Presentation to Test Attentional Blink for Nontargets
Published on: December 5, 2014
9.2K
Flexible Job Shop Scheduling via Dual Attention Network-Based Reinforcement Learning
IEEE Transactions on Neural Networks and Learning Systems
|September 11, 2023
Summary
This study introduces a new deep learning framework using attention models and reinforcement learning to solve complex flexible job shop scheduling problems (FJSP). The approach significantly improves scheduling efficiency and solution quality for manufacturing operations.
Area of Science:
- Operations Research
- Artificial Intelligence
- Manufacturing Systems Engineering
Background:
- Flexible job shop scheduling problems (FJSP) present significant complexity due to operations being processable on multiple machines.
- Existing deep reinforcement learning (DRL) methods for FJSP, while promising, often yield solutions with suboptimal quality compared to exact methods.
- There is a need for advanced methods to effectively capture intricate operation-machine relationships in FJSP.
Purpose of the Study:
- To develop a novel end-to-end learning framework for solving the flexible job shop scheduling problem (FJSP).
- To enhance decision-making in FJSP by integrating self-attention mechanisms with deep reinforcement learning.
- To improve the quality and scalability of solutions for complex manufacturing scheduling.
Main Methods:
- A dual-attention network (DAN) is proposed, incorporating interconnected operation and machine attention blocks for deep feature extraction.
- The DAN leverages self-attention to precisely model complex relationships between operations and machines.
- Deep reinforcement learning (DRL) is employed for scalable decision-making, guided by features extracted from the DAN.
Main Results:
- The proposed DAN-DRL framework significantly outperforms traditional priority dispatching rules (PDRs) and existing state-of-the-art DRL methods on synthetic and benchmark FJSP datasets.
- The approach achieves solution quality comparable to exact methods (e.g., OR-Tools) in specific scenarios.
- Demonstrated favorable generalization capabilities on large-scale and real-world unseen FJSP instances.
Conclusions:
- The novel end-to-end learning framework effectively addresses the limitations of current DRL approaches for FJSP.
- The dual-attention network successfully captures complex production dynamics, leading to high-quality scheduling decisions.
- This method offers a scalable and effective alternative for complex manufacturing scheduling, approaching the performance of exact methods.
Related Concept Videos
Reinforcement Schedules
202
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
202
Sequence Networks of Rotating Machines
123
A Y-connected synchronous generator, grounded through a neutral impedance, is designed to produce balanced internal phase voltages with only positive-sequence components. The generator's sequence networks include a source voltage that is exclusively in the positive-sequence network. The sequence components of line-to-ground voltages at the generator terminals illustrate this configuration.
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
123
Reinforcement
273
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
273
Observational Learning
207
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
207
Multi-input and Multi-variable systems
127
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
127
Avoidance Learning and Learned Helplessness
1.8K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.8K

