Related Experiment Video
Updated: Dec 24, 2025

An Open-Source Virtual Reality System for the Measurement of Spatial Learning in Head-Restrained Mice
Published on: March 3, 2023
Reinforcement Learning: Full Glass or Empty - Depends Who You Ask
Jacob J W Bakermans1, Timothy H Muller2, Timothy E J Behrens3
1Wellcome Centre for Integrative Neuroimaging, FMRIB, Nuffield Department of Clinical Neurosciences, University of Oxford, John Radcliffe Hospital, Oxford OX3 9DU, UK.
A new theory extends dopamine prediction error signaling by incorporating the full distribution of future rewards, not just the average. This approach better explains observed dopamine responses in the brain.
Area of Science:
- Neuroscience
- Computational Neuroscience
- Artificial Intelligence
Background:
- The prediction error theory of dopamine is a cornerstone in understanding reward-based learning.
- Existing models primarily focus on the average expected reward, which may not fully capture complex neural responses.
Purpose of the Study:
- To extend the prediction error theory of dopamine by incorporating the full distribution of future rewards.
- To investigate if this extended theory provides a more accurate account of dopamine responses compared to traditional models.
Main Methods:
- Importing concepts from artificial intelligence, specifically distributional reinforcement learning.
- Developing a computational model that represents the entire probability distribution of future rewards.
Main Results:
- The extended theory, by considering the full reward distribution, offers a more comprehensive explanation for dopamine neuron activity.
- This distributional approach better accounts for empirical observations of dopamine signaling under various reward contingencies.
Conclusions:
- Representing the full distribution of future rewards offers a more nuanced and accurate framework for understanding dopamine's role in learning and decision-making.
- This work highlights the potential of integrating advanced artificial intelligence concepts into neuroscience for improved theoretical models.
More Related Videos
13:24A Procedure to Observe Context-induced Renewal of Pavlovian-conditioned Alcohol-seeking Behavior in Rats
Published on: September 19, 2014
08:05A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Related Concept Videos
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Associative Learning
Classical conditioning, also known...
Purposive Learning