Related Experiment Video
Updated: Jul 15, 2026

09:43
A Fully Automated Rodent Conditioning Protocol for Sensorimotor Integration and Cognitive Control Experiments
Published on: April 15, 2014
Short-term memory traces for action bias in human reinforcement learning.
Rafal Bogacz1, Samuel M McClure, Jian Li
1Center for the Study of Brain, Mind and Behavior, Princeton University, Princeton, NJ 08544, USA. R.Bogacz@bristol.ac.uk
Brain Research
|April 27, 2007
Summary
Eligibility traces (ETs) improve reinforcement learning models. Our study shows human behavior in a sequential game is explained by temporal difference learning with ETs, suggesting a neural basis for these traces.
Area of Science:
- Neuroscience
- Computational Neuroscience
- Cognitive Science
Background:
- Reinforcement learning models how organisms learn from rewards and punishments.
- The credit assignment problem is central to reinforcement learning, concerning delayed rewards.
- Eligibility traces (ETs) enhance temporal difference learning by remembering past choices.
Purpose of the Study:
- To investigate if reinforcement learning models with long-lasting eligibility traces (ETs) can better explain human behavior.
- To examine the impact of timing between choices on human performance in a sequential economic decision game.
Main Methods:
- Human subjects played a sequential economic decision game with differing short-term and long-term optimal strategies.
- Behavioral data was analyzed to assess the influence of inter-choice intervals on performance.
- A temporal difference learning model incorporating ETs was used to explain observed human behavior.
Main Results:
- Human performance was significantly affected by the time between choices, in a counterintuitive manner.
- This behavioral pattern was accurately predicted by a temporal difference learning model with persistent ETs.
- Recent findings suggest short-term synaptic plasticity in dopamine neurons could mechanistically support these ETs.
Conclusions:
- Reinforcement learning models incorporating eligibility traces that persist across actions provide a compelling explanation for human decision-making in sequential tasks.
- The findings bridge computational models of learning with neurobiological mechanisms, highlighting dopamine neuron plasticity.
- This work advances our understanding of the neural underpinnings of learning and decision-making under delayed reward conditions.
Related Concept Videos
Timing and Consequences on Behavior
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant factor...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant factor...
Law of Effect
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Implicit Memories
Implicit memories, also known as non-declarative memories, are long-term memories that function outside of conscious awareness. These memories influence behavior and skills without explicit knowledge. This type of memory is evident in tasks like playing tennis, snowboarding, and texting. Implicit memory has three subsystems: procedural memory, conditioning, and priming. This type of memory is essential in various activities, from everyday tasks to specialized skills.
One key aspect of implicit...
One key aspect of implicit...
Purposive Learning
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a bonus...
Cognitive Learning
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Instinctive Drift
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
