Related Experiment Video
Updated: Mar 18, 2026

Combining Computer Game-Based Behavioural Experiments With High-Density EEG and Infrared Gaze Tracking
Published on: December 16, 2010
People Can Adaptively Exploit Model-free and Model-based Reinforcement Learning in Competitive Games
1Kansas State University, Manhattan, USA.
Abstract:
A key goal for organisms in competitive social interactions is learning strategies that outperform opponents. Despite the ample literature on modeling strategic behaviors in games, little research has parametrically examined the degree to which individuals' performance and strategy depend on an opponent's predictability and how they change over time. To test these questions, we conducted two experiments investigating people's behavior in the competitive game Rock, Paper, Scissors against computerized opponents programmed using reinforcement learning algorithms. In Experiment 1, the RL algorithms were model-free, where only the values of selected actions were updated following feedback. In Experiment 2, the algorithms were model-based, where the values of unselected actions were also updated. Results from both experiments showed that subjects significantly outperformed both classes of RL opponents, but they implemented different strategies to do so. Specifically, participants tended to engage in a Win-Stay/Lose-Shift strategy in Experiment 1 but a Win-Shift/Lose-Shift strategy in Experiment 2, contrary to behaviors predicted by typical RL models and learning theories. We discuss the theoretical and practical implications of shifting away from reinforced behaviors, reinforcement learning as a representative computational framework of strategic decision-making, and how future research can continue this investigation by testing additional models and competitive games.
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Reinforcement Schedules
Once a behavior is learned,...
Randomized Experiments
Simple randomization
Simple...