Related Experiment Video
Updated: Jul 15, 2026

Spatial Multiobjective Optimization of Agricultural Conservation Practices using a SWAT Model and an Evolutionary Algorithm
Published on: December 9, 2012
Evolution of strategies in iterated games: a deep multi-agent reinforcement learning approach with hybrid genetic and
ShangJun Li1, Avijit Deb Nath2, Md Hadi Al-Amin3
1School of Electrical Engineering and Computer Science, The University of Queensland, Brisbane, QLD, Australia.
Abstract:
The emergence of cooperation among self-interested agents in social dilemmas remains a central unsolved problem in evolutionary game theory, multi-agent systems, and behavioral science. While memory-one reciprocal strategies such as Tit-for-Tat have been extensively studied, the evolutionary properties of memory-two bilateral reciprocity strategies-which condition on a two-round bilateral history over a 216 = 65, 536-element strategy space-remain comparatively undercharacterised. This study presents a unified agent-based computational framework integrating a Q-learning multi-agent reinforcement learning (MARL) engine with Moran, Wright-Fisher, and replicator dynamics selection mechanisms to systematically investigate the evolutionary fate of memory-two (M2) bilateral reciprocity strategies across eight controlled simulation experiments. Under canonical Prisoner's Dilemma conditions (T = 5, R = 3, P = 1, S = 0, N = 100, μ = 0.05, β = 2.0, ε = 0.02), the simulation converged to a high-cooperation quasi-equilibrium with mean cooperation rate ρ C = 0.814 ± 0.029. The M2 strategy class collectively captured 61.0% of the evolutionary equilibrium population, with all three M2 strategies achieving fixation probability Pfix = 1.000 when invading an AllD-resident population-a 10-fold enrichment over the neutral expectation. Cooperation was robust to behavioral noise up to ε≈0.15, beyond which a noise-induced phase transition collapsed cooperative equilibria. Selection pressure, mutation rate, game class, and population size were systematically varied, confirming that M2 dominance is robust across all tested conditions. These results establish memory-two bilateral reciprocity as the dominant evolutionary strategy class in noisy iterated social dilemmas and provide a rigorous computational characterization of the conditions under which it emerges and persists.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Lagrange Multipliers: Problem Solving
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...