Related Experiment Video
Updated: Jan 11, 2026

The HoneyComb Paradigm for Research on Collective Human Behavior
Published on: January 19, 2019
A unified and efficient training framework for open-ended non-transitive games
Shaokang Dong1, Chao Li2, Shangdong Yang2
1School of Computer and Electronic Information, Nanjing Normal University, Nanjing, China; State Key Laboratory for Novel Software Technology, Nanjing University, China.
Abstract:
Self-play (SP) and Policy-Space Response Oracles (PSRO) are two fundamental training frameworks for solving Nash equilibrium (NE) in games. SP only performs effectively in transitive games, aiming to progressively derive a consistent winning strategy with an efficient warm start. In open-ended non-transitive games (e.g., Rock-Paper-Scissors), PSRO maintains a policy population and approximates the NE in the meta-game. However, PSRO requires training the policy from scratch in each iteration, making it inefficient in large-scale games. To address these limitations, we propose a Unified and Efficient Play (UEP) framework that combines the strengths of SP and PSRO, enabling the solution of NE in open-ended non-transitive games while benefiting from a warm start in each iteration. In addition, a unified regularized distance metric is proposed to trade off the accuracy of SP and the diversity of PSRO, enhancing the overall training efficiency. We present theoretical evidence that UEP can converge to the NE. Empirically, experiments are conducted in various open-ended games with strong non-transitivity. The results validate the superior performance of UEP in approximating NE and generating robust policies compared to prevailing SP and PSRO variants.
Related Concept Videos
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Observational Learning
Collisions in Multiple Dimensions: Introduction
Statically Indeterminate Problem Solving
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Multi-input and Multi-variable systems
In the absence of...