Related Experiment Video
Updated: Sep 9, 2025

A Conflict Model of Reward-seeking Behavior in Male Rats
Published on: February 20, 2019
How working memory and reinforcement learning interact when avoiding punishment and pursuing reward concurrently
Peter F Hitchcock1, Joonhwa Kim2, Michael J Frank3
1Department of Psychology, Emory University.
Abstract:
Humans learn adaptive behaviors via a durable but incremental reinforcement learning (RL) system and a fast but fleeting working memory (WM) system. Past work parsing these systems focused on reward learning alone; hence, little is known about how they interact while simultaneously learning to avoid punishment and whether arbitrating between these demands is disrupted by psychiatric symptoms. We administered a novel reward/punishment RL-WM task to an online sample oversampled for depression and anxiety symptoms (N = 298; n = 275 after quality control). Participants avoided punishment during initial learning, yet poorly retained this avoidance. Computational modeling captured this pattern via the fleeting WM system facilitating punishment avoidance, while the durable RL system retained little about punishment. Our task also included two test phases interleaved with learning, which permitted a targeted examination of past findings that WM blunts the RL system. When RL-based retention was tested midway through learning, we indeed found evidence of blunting. Yet, after learning resumed-leading to further prediction errors-blunting was no longer evident in a final test phase. However, individual differences moderated this effect: Some individuals were especially susceptible to blunting; for others, WM actually facilitated retention. Finally, task performance was largely spared as a function of depression/anxiety and trait rumination. Overall, our findings demonstrate that-when seeking to attain reward and avoid punishment concurrently-the WM system can facilitate short-term punishment avoidance while the RL system retains little about punishment, reveal individual differences in the extent to which WM blunts RL, and demonstrate intact behavior under internalizing-disorder symptoms. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
More Related Videos
14:24An Appetitive Spatial Working Memory Task for Mice in a Semi-Automated 8-Arm Radial Maze, Reducing Fearful Memory Association in the Maze
Published on: July 29, 2025
08:05A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Related Concept Videos
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Associative Learning
Classical conditioning, also known...
Working Memory
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...