Related Experiment Video
Updated: Jul 29, 2025

The HoneyComb Paradigm for Research on Collective Human Behavior
Published on: January 19, 2019
Optimal strategy of the simultaneous dice game Pig for multiplayers: when reinforcement learning meets game theory
Tian Zhu1, Merry Ma2, Lu Chen3
1Department of Applied Mathematics and Statistics, State University of New York at Stony Brook, Stony Brook, NY, 11794, USA. tian.zhu@stonybrook.edu.
Abstract:
In this work, we focus on using reinforcement learning and game theory to solve for the optimal strategies for the dice game Pig, in a novel simultaneous playing setting. First, we derived analytically the optimal strategy for the 2-player simultaneous game using dynamic programming, mixed-strategy Nash equilibrium. At the same time, we proposed a new Stackelberg value iteration framework to approximate the near-optimal pure strategy. Next, we developed the corresponding optimal strategy for the multiplayer independent strategy game numerically. Finally, we presented the Nash equilibrium for simultaneous Pig game with infinite number of players. To help promote the learning of and interest in reinforcement learning, game theory and statistics, we have further implemented a website where users can play both the sequential and simultaneous Pig game against the optimal strategies derived in this work.
More Related Videos
07:59Transition of Farm Pigs to Research Pigs using a Designated Checklist followed by Initiation of Clicker Training - a Refinement Initiative
Published on: August 21, 2021
06:18The Collective Trust Game: An Online Group Adaptation of the Trust Game Based on the HoneyComb Paradigm
Published on: October 20, 2022
Related Concept Videos
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...