Related Experiment Video
Updated: Jun 6, 2026

Peering into the Dynamics of Social Interactions: Measuring Play Fighting in Rats
Published on: January 18, 2013
Population-dependent agent performance in non-transitive games: a multi-agent rock-paper-scissors benchmark
Ou Deng1, Jianting Xu2, Shoji Nishimura3
1Graduate School of Human Sciences, Waseda University, Tokorozawa, Saitama, 359-1192, Japan. dengou@toki.waseda.jp.
None:
Non-transitive environments complicate the notion of a single "best" strategy: performance depends on the opponent population, and rankings are meaningful only relative to a specified opponent pool and protocol. We present a reproducible multi-agent benchmark for iterated rock-paper-scissors and evaluate 54 agents from 18 archetypes-deep recurrent and transformer sequence models, actor-critic reinforcement learners, Bayesian/Markov predictors, classical classifiers, and rule-based baselines-in 500-round double round-robin tournaments across 10 random seeds. To support auditability, we define a simple regret certificate: a Lipschitz-type inequality that upper-bounds a best-response payoff gap using the [Formula: see text] discrepancy between an agent's predicted action distribution and an empirical estimate of the opponent's action distribution, computable online from logged predictions. Our experiments indicate that (i) recurrent predictors tend to achieve the highest and most stable scores, with gains that are largest against predictable opponents; (ii) rankings shift notably with the opponent pool (Spearman [Formula: see text] between two evaluation rosters), with the top-ranked method changing across configurations; and (iii) the induced meta-game exhibits substantial non-transitivity, including 134 detected three-cycles in the pairwise payoff matrix. Under our 500-round online update budget and short-context design, transformer agents are competitive but do not outperform tuned recurrent baselines, which may reflect an inductive-bias mismatch in short-horizon adversarial play. Our code and analysis pipeline provide an extensible testbed for studying population-dependent evaluation and learning dynamics in canonical non-transitive games.
Related Concept Videos
Inclusive Fitness
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Social Facilitation
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can have a...