Related Experiment Video
Updated: Sep 3, 2025

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
Reward prediction errors, not sensory prediction errors, play a major role in model selection in human reinforcement
Yihao Wu1, Masahiko Morita2, Jun Izawa2
1School of Integrative and Global Majors, University of Tsukuba, Tennodai 1-1-1, Tsukuba, Ibaraki, 305-8573, Japan.
The human brain appears to use reward prediction errors, not sensory prediction errors, to select appropriate internal models during reinforcement learning. This finding advances our understanding of cognitive flexibility and decision-making processes.
Area of Science:
- Neuroscience
- Cognitive Science
- Computational Neuroscience
Background:
- Model-based reinforcement learning (RL) is crucial for adapting to dynamic environments.
- The brain's mechanism for selecting appropriate internal models during RL remains poorly understood.
- Current theories propose either sensory prediction errors or reward prediction errors drive model selection.
Purpose of the Study:
- To investigate the neural mechanisms of internal model selection in human reinforcement learning.
- To compare the predictive power of sensory prediction error-based versus reward prediction error-based models of human behavior.
- To determine whether the brain utilizes sensory or reward prediction errors for selecting environmental models.
Main Methods:
- Designed a switching experiment transitioning between first-order and second-order Markov decision processes.
- Tested two computational models: a sensory prediction-error-driven Bayesian algorithm and a reward-prediction-error-driven policy gradient algorithm.
- Compared model simulations against human behavioral data from the RL task.
Main Results:
- Model fitting analyses indicated that the policy gradient algorithm provided a better fit to human RL behavior than the Bayesian algorithm.
- This suggests that reward prediction errors play a more significant role in model selection than sensory prediction errors.
- Human participants' behavior aligned more closely with predictions from the reward-prediction-error-driven model.
Conclusions:
- The human brain likely employs reward prediction errors to select appropriate internal models during reinforcement learning.
- This finding challenges the primacy of sensory prediction errors in internal model selection.
- The study provides evidence for a reward-based mechanism underlying cognitive flexibility in RL tasks.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Reinforcement Schedules
Once a behavior is learned,...
Primary and Secondary Reinforcers
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...

