Related Experiment Video
Updated: Aug 22, 2025

Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
Value-free reinforcement learning: policy optimization as a minimal model of operant behavior
Daniel Bennett1,2, Yael Niv1,3, Angela J Langdon1
1Princeton Neuroscience Institute, Princeton University, USA.
Policy-gradient reinforcement learning models offer a simpler alternative to value-based models for understanding decision-making. These models optimize behavior directly, providing a compelling explanation for operant behavior in neuroscience and economics.
Area of Science:
- Cognitive Neuroscience
- Neuroeconomics
- Computational Psychiatry
Background:
- Reinforcement learning (RL) is crucial for modeling learning and decision-making.
- Current research predominantly uses value-based RL models, assuming decisions are based on comparing action values.
- An alternative, policy-gradient RL, optimizes behavior directly without explicit value-learning.
Purpose of the Study:
- To review behavioral and neural evidence.
- To demonstrate that policy-gradient models offer a more parsimonious explanation for observed findings compared to value-based models.
- To highlight the utility of policy-gradient models in understanding operant behavior.
Main Methods:
- Literature review of recent behavioral and neural findings.
- Comparative analysis of explanatory power between value-based and policy-gradient RL models.
- Synthesis of evidence supporting policy-gradient approaches.
Main Results:
- Several behavioral and neural findings are more simply explained by policy-gradient models.
- Policy-gradient models bypass the need for intermediate value-learning steps.
- Direct policy optimization provides a compelling account of operant behavior.
Conclusions:
- Policy-gradient reinforcement learning presents a lightweight and effective alternative to value-based models.
- These models offer a valuable framework for understanding decision-making processes.
- The study advocates for broader consideration of policy-gradient models in cognitive neuroscience and neuroeconomics.
Related Concept Videos
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Operant Conditioning Intervention
In operant conditioning, behaviors that are...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Behaviorism
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...
Associative Learning
Classical conditioning, also known...

