Related Experiment Video
Updated: Jun 9, 2026

New Variations for Strategy Set-shifting in the Rat
Published on: January 23, 2017
Alterations in choice behavior by manipulations of world model
Abstract:
How to compute initially unknown reward values makes up one of the key problems in reinforcement learning theory, with two basic approaches being used. Model-free algorithms rely on the accumulation of substantial amounts of experience to compute the value of actions, whereas in model-based learning, the agent seeks to learn the generative process for outcomes from which the value of actions can be predicted. Here we show that (i) "probability matching"-a consistent example of suboptimal choice behavior seen in humans-occurs in an optimal Bayesian model-based learner using a max decision rule that is initialized with ecologically plausible, but incorrect beliefs about the generative process for outcomes and (ii) human behavior can be strongly and predictably altered by the presence of cues suggestive of various generative processes, despite statistically identical outcome generation. These results suggest human decision making is rational and model based and not consistent with model-free learning.
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Behavior Modification
A real-world application of operant conditioning principles is applied...
Impression Management Techniques IV: Altercasting
Counterfactual Thinking
Schemas
Operant Conditioning Intervention
In operant conditioning, behaviors that are...
