Related Experiment Video
Updated: May 21, 2025

A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Deep reinforcement learning can promote sustainable human behaviour in a common-pool resource problem
Raphael Koster1, Miruna Pîslar2, Andrea Tacchetti3
1Google DeepMind, London, UK. rkoster@google.com.
Deep reinforcement learning designed a social planner for sustainable resource sharing in trust games. This AI mechanism promoted fairer and larger collective returns by adapting generosity and sanctioning defectors.
Area of Science:
- Behavioral Economics
- Artificial Intelligence
- Game Theory
Background:
- Social dilemmas, such as resource allocation, often involve a conflict between individual gain and collective well-being.
- Effective resource allocation mechanisms are crucial for sustaining shared resources and promoting cooperation.
- Previous mechanisms often fail to adapt to dynamic player behavior and resource availability.
Purpose of the Study:
- To design an artificial intelligence (AI) based social planner using deep reinforcement learning (RL) to promote sustainable contributions in a multiplayer trust game.
- To create a simulated economy of human-like players to test and refine the RL mechanism.
- To maximize aggregate returns for all players while fostering fairness and cooperation.
Main Methods:
- Trained neural networks to simulate human player behavior in an iterated multiplayer trust game.
- Employed deep reinforcement learning (RL) to train a social planner mechanism.
- Developed a redistributive policy that conditions generosity on resource availability and incorporates temporary sanctions for defectors.
Main Results:
- The RL-designed mechanism significantly increased aggregate returns compared to baseline mechanisms.
- The mechanism achieved a more equal distribution of the surplus among players.
- The AI's policy, when translated into an explainable mechanism, was more favorably received by human participants.
Conclusions:
- Deep reinforcement learning can effectively design adaptive mechanisms for social dilemmas, promoting sustainable cooperation.
- AI-driven resource allocation can lead to both increased efficiency and greater equity in shared resource management.
- Explainable AI-derived policies enhance player acceptance and participation in cooperative systems.
Related Concept Videos
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Primary and Secondary Reinforcers
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Instinctive Drift
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Operant Conditioning Intervention
In operant conditioning, behaviors that are...

