Related Experiment Video
Updated: Aug 6, 2026

A Conflict Model of Reward-seeking Behavior in Male Rats
Published on: February 20, 2019
Competing value signals impair reward-learning via dopaminergic mechanisms and increase exploration
Wen-Wei Lin1, Pei-Yu Lee1, Hsin-Yun Tsai2
1Graduate Institute of Brain and Mind Sciences, National Taiwan University College of Medicine, Taipei, Taiwan.
None:
Effective reinforcement learning requires balancing exploration of uncertain options with exploitation of known outcomes. In real-world contexts, the same action may yield rewards in some situations and punishments in others, yet how these learning processes influence each other remains unclear. Here, we examine the neural mechanisms underlying how reward and punishment learning interact to guide adaptive behavior. We conducted four experiments (N = 159) using an instrumental learning task with binary choices, some of which were exclusive to reward or punishment learning trials, while others appeared in both, allowing assessment of their interaction. When choices were tied to a single learning type, reward learning engages less exploration (i.e., fewer choices of lower-value options) than punishment learning. Critically, when both learning processes were concurrently engaged, reward learning was selectively impaired, accompanied by enhanced exploration and greater activation in exploration-related prefrontal regions revealed by fMRI. Computational modeling showed that impaired reward learning was best explained by sensitivity to prior punishment history associated with reward-learning options, while individual differences in loss aversion predicted the degree of increased exploration. Finally, pharmacological attenuation of dopaminergic signaling via the D2/3 receptor antagonist amisulpride abolished both the increased exploration and the interference with reward learning. These findings suggest that punishment-history interference during reward learning is dopamine-modulated and associated with increased exploration, with individual differences in this exploration linked to loss aversion, providing a mechanistic account of how the brain resolves competing value signals and informing dopamine-related learning disturbances in neuropsychiatric conditions.
Related Concept Videos
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant factor...
Incentive Theory: Pull Theory of Motivation
The theory differentiates between intrinsic and...
Adrenergic Agonists: Indirect-Acting Agents
One mechanism involves depleting stored catecholamines by displacing them from synaptic vesicles. These agents, known as "displacers," are transported into vesicles at the expense of noradrenaline. Examples include amphetamine and tyramine, which lack a catechol moiety, resulting in prolonged action, improved oral bioavailability, and...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Primary and Secondary Reinforcers
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...

