Related Experiment Video
Updated: Oct 4, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.1K
Outracing champion Gran Turismo drivers with deep reinforcement learning
Peter R Wurman1, Samuel Barrett2, Kenta Kawamoto3
1Sony AI, New York, NY, USA. peter.wurman@sony.com.
Nature
|February 10, 2022
Summary
Researchers developed artificial intelligence (AI) agents for the Gran Turismo racing simulation that can compete with top human e-sports drivers. This AI combines speed and tactics, demonstrating advanced control in complex, real-time environments.
Area of Science:
- Artificial Intelligence
- Robotics
- Computational Neuroscience
Background:
- Real-time decision-making in physical systems interacting with humans presents significant challenges.
- Automobile racing exemplifies these challenges, requiring complex tactical maneuvers and precise vehicle control at physical limits.
- Racing simulations like Gran Turismo accurately model these complex, multi-agent dynamics.
Purpose of the Study:
- To train artificial intelligence (AI) agents capable of competing at the highest level in the Gran Turismo racing simulation.
- To develop an integrated control policy combining speed and tactical decision-making for AI racers.
- To create a reward function that promotes competitive yet sportsmanlike behavior within the simulation's rules.
Main Methods:
- Utilized state-of-the-art, model-free, deep reinforcement learning algorithms.
- Implemented mixed-scenario training to expose agents to diverse racing conditions.
- Developed a specialized reward function to balance performance with adherence to sportsmanship norms.
Main Results:
- Trained AI agents, named Gran Turismo Sophy, demonstrated the ability to compete with world-class e-sports drivers.
- The AI agents exhibited exceptional speed and sophisticated tactical maneuvers.
- Gran Turismo Sophy successfully won a head-to-head competition against four top human Gran Turismo players.
Conclusions:
- The study showcases the potential of deep reinforcement learning for controlling complex dynamical systems in human-interactive domains.
- AI agents can achieve championship-level performance in simulated environments requiring both speed and strategic interaction.
- Controlling AI in domains with imprecise human norms, like sportsmanship, presents both possibilities and challenges.
Related Concept Videos
Reinforcement
411
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
411
Hierarchy of Motor Control
3.9K
The hierarchy of motor control refers to the different levels of organization and processing involved in controlling movement in the body. These levels range from higher cortical areas involved in planning and decision-making to lower spinal cord reflexes that respond automatically to external stimuli.
3.9K
Rolling Resistance: Problem Solving
486
Rolling resistance, also known as rolling friction, is the force that resists the motion of a rolling object, such as a wheel, tire, or ball, when it moves over a surface. It is caused by the deformation of the object and the surface in contact with each other, as well as other factors like internal friction, hysteresis, and energy losses within the materials. Rolling resistance opposes the object's motion, requiring additional energy to overcome it and maintain movement. In practical...
486
Torque
16.0K
Torque is an important quantity for describing the dynamics of a rotating rigid body. We see the application of torque in many ways in the world, such as when pressing the accelerator in a car, which causes the engine to apply additional torque on the drivetrain. Here, we define torque and provide a framework to create an equation to calculate torque for a rigid body with fixed-axis rotation.
Torque can be considered as the rotational counterpart to force. Since forces change the translational...
Torque can be considered as the rotational counterpart to force. Since forces change the translational...
16.0K
Observational Learning
360
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
360
End Point Prediction: Gran Plot
668
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
668

