Related Experiment Video
Updated: Feb 1, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play.
David Silver1,2, Thomas Hubert3, Julian Schrittwieser3
1DeepMind, 6 Pancras Square, London N1C 4AG, UK. davidsilver@google.com dhcontact@google.com.
A new artificial intelligence algorithm, AlphaZero, uses reinforcement learning to achieve superhuman performance in games like chess and Go. It learns from self-play without human domain knowledge, defeating top programs from random play.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Game Theory
Background:
- Chess has been a long-standing benchmark for artificial intelligence (AI) research.
- Current AI chess programs rely on human expertise and handcrafted evaluation functions.
- Recent advances in reinforcement learning (RL) have shown promise in complex games like Go.
Purpose of the Study:
- To develop a generalized AI algorithm capable of achieving superhuman performance across multiple challenging games.
- To demonstrate the efficacy of reinforcement learning from self-play without domain-specific knowledge.
Main Methods:
- The study introduces the AlphaZero algorithm, a unified approach generalizing the AlphaGo Zero methodology.
- AlphaZero utilizes deep neural networks and Monte Carlo tree search, trained via reinforcement learning from self-play.
- The algorithm is provided only with the rules of the game, starting from random play.
Main Results:
- AlphaZero achieved superhuman performance in chess, shogi, and Go.
- The algorithm convincingly defeated world-champion level AI programs in these games.
- This demonstrates a significant advancement in generalized game-playing AI.
Conclusions:
- Reinforcement learning from self-play is a powerful paradigm for developing superhuman AI in complex strategic domains.
- AlphaZero represents a significant step towards more general artificial intelligence systems.
- The approach removes the need for extensive domain-specific engineering in AI game development.
Related Concept Videos
Master Transcription Regulators
Master Transcription Regulators
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Social Foundations of Self I: Play and Game
Corrosion of Reinforcement
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
Reinforcement Schedules
Once a behavior is learned,...

