Related Experiment Video
Updated: May 14, 2026

06:57
Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
The Reward Positivity Tracks Positive Reward Prediction Errors From Feedback to Cues During Reinforcement Learning.
Yifan Gao1, Robert Wilson2, Galit Karpov1,3
1Center for Molecular and Behavioral Neuroscience, Rutgers University, Newark, New Jersey, USA.
Psychophysiology
|May 13, 2026
Summary
The brain learns to predict rewards by shifting reward prediction errors (RPEs) from outcomes to cues. This study shows the reward positivity signal, reflecting RPEs, transfers from feedback to cues during learning.
Area of Science:
- Neuroscience
- Cognitive Science
- Computational Neuroscience
Background:
- Temporal difference (TD) learning theory predicts that reward prediction errors (RPEs) shift temporally during learning.
- The reward positivity (RewP) is an electrophysiological signal linked to anterior midcingulate cortex sensitivity to positive RPEs.
- It's hypothesized that RewP should transfer from outcome feedback to predictive cues as learning progresses, but this remains largely untested.
Purpose of the Study:
- To investigate the temporal dynamics of the reward positivity (RewP) during reinforcement learning.
- To test the core prediction of TD learning theory regarding the temporal backpropagation of RPEs.
- To examine individual differences in learning rates and their relationship with RewP shifts.
Main Methods:
- Electroencephalography (EEG) was recorded from 73 healthy adults during an extended probabilistic selection task (PST).
- RewP amplitude was measured at cue and feedback during early and late learning phases.
- Participants were categorized into rapid and slow learners, and Q-learning models were used to estimate learning rates.
Main Results:
- Evidence of temporal backpropagation was observed: RewP shifted from feedback to predictive cues as learning progressed.
- Rapid learners exhibited more pronounced RewP shifts, correlating with higher learning rates.
- The RewP disappeared from feedback responses late in learning, emerging instead at cue presentation.
Conclusions:
- This study provides the first clear evidence for temporal backpropagation of the reward positivity, supporting TD learning theory.
- The findings validate RewP as a neural marker for positive RPEs.
- Understanding these neural dynamics is crucial for characterizing reinforcement learning and individual differences.
Related Concept Videos
Reinforcement
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Operant Conditioning
Operant conditioning, a key concept in behavioral psychology, involves using reinforcement and punishment to alter the likelihood of a behavior being repeated. B.F. introduced this type of conditioning. Skinner focused on voluntary behaviors and the consequences that follow them, influencing whether these behaviors will be strengthened or diminished.
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Reinforcement Schedules
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
Law of Effect
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Primary and Secondary Reinforcers
In psychology, reinforcement is a key concept in behavior modification. B.F. Skinner demonstrated this with his experiments involving rats in what is known as a Skinner box. The rats learned to press a lever to receive food, a primary reinforcer that fulfilled their innate need for nourishment.
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
Incentive Theory: Pull Theory of Motivation
Incentive theory, or the "pull theory" of motivation, suggests that external rewards primarily drive behavior. Individuals are motivated to engage in activities when they anticipate a desirable outcome. This is why people often work hard for promotions or study intensively to achieve high grades. These incentives can be tangible, physical rewards such as money or promotions, or intangible, non-physical rewards like praise and social recognition.
The theory differentiates between intrinsic and...
The theory differentiates between intrinsic and...

