Related Experiment Video
Updated: Jan 4, 2026

06:04
Study Motor Skill Learning by Single-pellet Reaching Tasks in Mice
Published on: March 4, 2014
22.0K
Grandmaster level in StarCraft II using multi-agent reinforcement learning.
Oriol Vinyals1, Igor Babuschkin2, Wojciech M Czarnecki2
1DeepMind, London, UK. vinyals@google.com.
Nature
|November 1, 2019
Summary
AlphaStar, an artificial intelligence agent, achieved Grandmaster level in StarCraft II by using multi-agent reinforcement learning. This AI demonstrates advanced capabilities in complex real-world strategy games.
Area of Science:
- Artificial Intelligence
- Computational Game Theory
- Multi-Agent Systems
Background:
- StarCraft presents complex multi-agent challenges relevant to real-world applications.
- Previous AI agents in StarCraft have relied on game simplification or superhuman abilities.
- No prior AI has matched the overall skill of top human StarCraft players.
Purpose of the Study:
- To develop an AI agent capable of competing at a top human level in the complex real-time strategy game StarCraft II.
- To utilize general-purpose learning methods applicable to other challenging domains.
Main Methods:
- Employed a multi-agent reinforcement learning algorithm.
- Utilized a diverse league of deep neural networks with adaptive strategies and counter-strategies.
- Trained on data from both human and agent games in StarCraft II.
Main Results:
- The AlphaStar agent achieved Grandmaster level across all three StarCraft II races.
- AlphaStar's performance surpassed 99.8% of ranked human players.
- Demonstrated a significant advancement in AI capabilities for complex strategy games.
Conclusions:
- General-purpose learning methods, particularly multi-agent reinforcement learning, can achieve human-level performance in complex strategic environments like StarCraft II.
- AlphaStar represents a major milestone in artificial intelligence research for real-time strategy games.
- The approach is extensible to other domains requiring complex decision-making and coordination.
Related Concept Videos
Reinforcement
765
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
765
Observational Learning
769
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
769
Reinforcement Schedules
414
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
414
Avoidance Learning and Learned Helplessness
2.4K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.4K
Multi-input and Multi-variable systems
358
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
358

