Related Experiment Videos
Reinforcement learning with reputation-based adaptive exploration promotes cooperation
An Li1,2, Wenqiang Zhu2,3,4, Chaoqian Wang5
1School of Mathematical Sciences, Beihang University, Beijing 100191, China.
Chaos (Woodbury, N.Y.)
|August 7, 2026
Summary
Agents in social dilemmas can learn to cooperate by adjusting their exploration rates based on reputation. Combining adaptive exploration with asymmetric reputation updating significantly boosts cooperation, especially for low-reputation agents.
Area of Science:
- Computational Social Science
- Behavioral Economics
- Artificial Intelligence
Background:
- Reinforcement learning models behavior adjustment via feedback.
- Q-learning uses exploration rates, typically constant, to balance action selection.
- Social evaluation systems imply exploration costs/benefits vary with reputation.
Purpose of the Study:
- To develop a spatial prisoner's dilemma model where Q-learning agents adapt exploration rates based on local reputation.
- To investigate how asymmetric, state-dependent reputation updating influences cooperation.
- To analyze the combined effects of adaptive exploration and reputation dynamics on cooperation.
Main Methods:
- Developed a spatial prisoner's dilemma model with Q-learning agents.
- Implemented adaptive exploration rates dependent on local reputation differences.
- Utilized an asymmetric, state-dependent rule for reputation updates.
- Analyzed spatial organization and stability of cooperation.
Main Results:
- Adaptive exploration and asymmetric reputation updating individually promote cooperation.
- Their combination yields a stronger cooperative effect than either mechanism alone.
- Low-reputation agents increase exploration to recover reputation; high-reputation agents decrease it to avoid losses.
- A stable checkerboard pattern of cooperation emerges at intermediate reputation concern.
- Stronger asymmetric reputation updating mitigates cooperation disruption caused by intermediate exploration rates.
Conclusions:
- Reputation dynamically regulates exploratory behavior during learning, stabilizing cooperation.
- Adaptive exploration based on reputation enhances cooperation more than fixed exploration rates.
- The model demonstrates how social reputation can foster prosocial behavior in learning agents.
Related Concept Videos
Reinforcement
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Avoidance Learning and Learned Helplessness
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Reinforcement Schedules
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
Social Exchange Theory
We have discussed why we form relationships, what attracts us to others, and different types of love. But what determines whether we are satisfied with and stay in a relationship? One theory that provides an explanation is social exchange theory. According to social exchange theory, we act as naïve economists in keeping a tally of the ratio of costs and benefits of forming and maintaining a relationship with others (Rusbult & Van Lange, 2003).
Social Exchange Theory
As formulated by John Thibaut and Harold Kelley, Social Exchange Theory explains human relationships as economic-like exchanges that maximize rewards and minimize costs. This theory suggests that individuals engage in relationships to gain benefits and reduce burdens, similar to economic transactions. It has been widely applied to various types of relationships, including romantic, professional, and social interactions.Rewards and Costs in RelationshipsRelationship rewards include emotional...