Related Experiment Video
Updated: Apr 18, 2026

Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
Modulation of feature attention by reward prediction error explains value learning behavior
Mingze L Leukos1, Albert Liang1, Grace W Lindsay1,2
1Department of Psychology, New York University, New York, NY 10003, USA.
None:
Adaptive behavior requires learning the value of environmental features while selectively attending to those most likely to yield reward. Reward prediction errors (RPEs) drive value learning and learned values guide attention, yet the computational function linking RPEs to attentional modulation remains unspecified. Here, we developed a reinforcement learning model with a perceptual front-end to investigate how value and RPE signals modulate attentional gain during learning. We compared five candidate RPE-attention transfer functions, each combined with either single- or multi-focus attention, against behavioral data from two adult male rhesus macaques performing a color-value learning task with shifting reward contingencies. Monkeys exhibited rapid initial learning followed by sub-optimal asymptotic accuracy. Overall, single-focus architectures consistently outperformed multi-focus counterparts on matching monkey errors, indicating that macaques collapse the value distribution into a winner-take-all attentional focus. Furthermore, the "Switch" model, in which attention targets the highest-valued feature but transiently inverts following negative RPEs, produced the fastest exploration dynamics following target switches and, together with the Absolute Value model, yielded decision confidence trajectories that positively correlated with empirical reaction times. In support of this, single-neuron correlation analyses revealed that 27 - 42% of neurons in prefrontal cortex, frontal eye fields, and lateral intraparietal area encoded previous-trial RPE at the time of next trial onset. In total, we conclude that capacity constrained attention that inverts its focus after negative RPE best explains value learning dynamics. These results provide a normative account for why biological learners sacrifice asymptotic precision for rapid adaptation in volatile environments.
Related Concept Videos
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Purposive Learning
Behaviorism
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...
Incentive Theory: Pull Theory of Motivation
The theory differentiates between...
Observational Learning

