Related Experiment Video
Updated: Jul 16, 2026

05:41
A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
Biological implementation of the temporal difference algorithm for reinforcement learning: theoretical comment on
1Department of Physiology, Institute for Neuroscience, Feinberg School of Medicine, Northwestern University, Chicago, IL 60611, USA. j-houk@northwestern.edu
Behavioral Neuroscience
|February 28, 2007
Summary
This study evaluates a new reinforcement learning model, Primary Value and Learned Value (PVLV), for brain function. The author proposes refining the existing Temporal Difference (TD) algorithm instead of adopting the PVLV model.
Area of Science:
- Neuroscience
- Computational Neuroscience
- Reinforcement Learning
Background:
- Brain's predictive capacity is crucial for survival, detecting reward/punishment predictors.
- Temporal Difference (TD) algorithm is the dominant model for reinforcement learning in neuroscience.
- A new model, Primary Value and Learned Value (PVLV), is proposed as simpler and more biologically realistic than TD.
Purpose of the Study:
- To evaluate the proposed Primary Value and Learned Value (PVLV) model.
- To compare the PVLV model with the existing Temporal Difference (TD) algorithm.
- To suggest an alternative approach to modeling predictive learning in the brain.
Main Methods:
- Commentary on a proposed new model (PVLV).
- Comparison of PVLV with existing Temporal Difference (TD) algorithm.
- Discussion of biological realism and model simplicity.
Main Results:
- The author argues against adopting the PVLV model.
- The author suggests modifications to a previous biological implementation of the TD algorithm.
- The commentary favors refining the TD algorithm over the new PVLV model.
Conclusions:
- The existing Temporal Difference (TD) algorithm, with modifications, remains a viable framework.
- The proposed PVLV model may not be necessary or superior to refined TD models.
- Further research could focus on biologically implementing and refining TD algorithms for predictive learning.
Related Concept Videos
Law of Effect
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Observational Learning
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning because...
Reinforcement Schedules
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...