Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reinforcement Schedules01:24

Reinforcement Schedules

250
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
250
Reinforcement01:23

Reinforcement

407
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
407
Observational Learning01:12

Observational Learning

354
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
354
Time-Domain Interpretation of PD Control01:07

Time-Domain Interpretation of PD Control

190
Proportional-Derivative (PD) control is a widely used control method in various engineering systems to enhance stability and performance. In a system with only proportional control, common issues include high maximum overshoot and oscillation, observed in both the error signal and its rate of change. This behavior can be divided into three distinct phases: initial overshoot, subsequent undershoot, and gradual stabilization.
Consider the example of control of motor torque. Initially, a positive...
190
Decision Making: P-value Method01:09

Decision Making: P-value Method

5.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.8K
Neural Regulation01:37

Neural Regulation

40.4K
Digestion begins with a cephalic phase that prepares the digestive system to receive food. When our brain processes visual or olfactory information about food, it triggers impulses in the cranial nerves innervating the salivary glands and stomach to prepare for food.
40.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

An abstract relational map emerges in the human medial prefrontal cortex with consolidation.

Current biology : CB·2026
Same author

Neural signatures of model-based and model-free reinforcement learning across prefrontal cortex and striatum.

eLife·2026
Same author

Human hippocampal ripples align new experiences with a grid-like schema.

Neuron·2025
Same author

A cognitive map for value-guided choice in the ventromedial prefrontal cortex.

Cell·2025
Same author

Neural mechanisms of credit assignment for delayed outcomes during contingent learning.

eLife·2025
Same author

Constructing future behavior in the hippocampal formation through composition and replay.

Nature neuroscience·2025

Related Experiment Video

Updated: Sep 30, 2025

A Fully Automated and Highly Versatile System for Testing Multi-cognitive Functions and Recording Neuronal Activities in Rodents
09:13

A Fully Automated and Highly Versatile System for Testing Multi-cognitive Functions and Recording Neuronal Activities in Rodents

Published on: May 3, 2012

14.5K

Reinforcement learning: Dopamine ramps with fuzzy value estimates.

James C R Whittington1, Timothy E J Behrens2

  • 1Wellcome Centre for Integrative Neuroimaging, University of Oxford, Oxford OX3 9DU, UK.

Current Biology : CB
|March 15, 2022
PubMed
Summary

A new reinforcement learning theory explains dopamine neuron behavior. The study extends temporal difference algorithms to unbiased learning, resolving state uncertainty and neuron ramping.

More Related Videos

Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

11.1K
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

8.8K

Related Experiment Videos

Last Updated: Sep 30, 2025

A Fully Automated and Highly Versatile System for Testing Multi-cognitive Functions and Recording Neuronal Activities in Rodents
09:13

A Fully Automated and Highly Versatile System for Testing Multi-cognitive Functions and Recording Neuronal Activities in Rodents

Published on: May 3, 2012

14.5K
Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

11.1K
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
07:52

Investigating Motor Skill Learning Processes with a Robotic Manipulandum

Published on: February 12, 2017

8.8K

Area of Science:

  • Neuroscience
  • Machine Learning Theory

Background:

  • Dopamine neurons exhibit ramping activity during decision-making tasks.
  • Existing reinforcement learning models do not fully account for this behavior under uncertainty.

Purpose of the Study:

  • To develop a theoretical framework explaining dopamine neuron ramping.
  • To investigate the role of state uncertainty in reinforcement learning.

Main Methods:

  • Extending the temporal difference algorithm for unbiased learning.
  • Mathematical modeling of reinforcement learning under state uncertainty.

Main Results:

  • The extended temporal difference algorithm successfully explains dopamine neuron ramping.
  • Unbiased learning under state uncertainty is key to this explanation.

Conclusions:

  • Reinforcement learning theory provides a mechanism for understanding neural computations.
  • This work bridges computational neuroscience and machine learning.