Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

88
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
88

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Orbitofrontal noradrenaline supports adaptive learning-rate adjustment in probabilistic reversal learning.

Proceedings of the National Academy of Sciences of the United States of America·2026
Same author

Dopaminergic mechanisms of dynamical social specialization.

Nature·2026
Same author

Why learning progress needs absolute values: Comment on Poli et al. (2024).

The European journal of neuroscience·2024
Same author

Strong and weak alignment of large language models with human values.

Scientific reports·2024
Same author

Rat anterior cingulate neurons responsive to rule or strategy changes are modulated by the hippocampal theta rhythm and sharp-wave ripples.

The European journal of neuroscience·2024
Same author

Adaptive Responding to Stimulus-Outcome Associations Requires Noradrenergic Transmission in the Medial Prefrontal Cortex.

The Journal of neuroscience : the official journal of the Society for Neuroscience·2024

Related Experiment Video

Updated: Jun 23, 2025

A Fully Automated Rodent Conditioning Protocol for Sensorimotor Integration and Cognitive Control Experiments
09:43

A Fully Automated Rodent Conditioning Protocol for Sensorimotor Integration and Cognitive Control Experiments

Published on: April 15, 2014

10.6K

Regulation of reinforcement learning parameters captures long-term changes in rat behaviour.

François Cinotti1,2, Etienne Coutureau3, Mehdi Khamassi1

  • 1Institut des Systèmes Intelligents et de Robotique, Sorbonne Université, CNRS, Paris, France.

The European Journal of Neuroscience
|June 26, 2024
PubMed
Summary

Rats progressively stabilized their learning strategies during pretraining on a decision-making task. A meta-learning model explained how they adjusted learning rate or inverse temperature based on average reward rate.

Keywords:
decision‐makingdopamineexploration‐exploitation trade‐offmeta‐learning

More Related Videos

Operant Procedures for Assessing Behavioral Flexibility in Rats
08:30

Operant Procedures for Assessing Behavioral Flexibility in Rats

Published on: February 15, 2015

20.8K
Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

10.9K

Related Experiment Videos

Last Updated: Jun 23, 2025

A Fully Automated Rodent Conditioning Protocol for Sensorimotor Integration and Cognitive Control Experiments
09:43

A Fully Automated Rodent Conditioning Protocol for Sensorimotor Integration and Cognitive Control Experiments

Published on: April 15, 2014

10.6K
Operant Procedures for Assessing Behavioral Flexibility in Rats
08:30

Operant Procedures for Assessing Behavioral Flexibility in Rats

Published on: February 15, 2015

20.8K
Pavlovian Conditioned Approach Training in Rats
06:57

Pavlovian Conditioned Approach Training in Rats

Published on: February 4, 2016

10.9K

Area of Science:

  • Neuroscience
  • Computational Biology
  • Animal Behavior

Background:

  • Animals in unpredictable environments face a learning dilemma: exploit current knowledge or explore for new information.
  • The regulation of this explore-exploit trade-off during initial learning remains poorly understood.

Purpose of the Study:

  • To investigate how rats adapt their reinforcement learning strategies over time during task acquisition.
  • To compare computational models for explaining long-term behavioral changes in learning.

Main Methods:

  • Observed 24 rats over 24 days on a three-armed bandit task.
  • Analyzed daily changes in rat performance and win-shift tendency.
  • Utilized computational modeling, including meta-learning approaches.

Main Results:

  • Rat performance and win-shift tendencies showed progressive stabilization across pretraining days.
  • Behavioral adaptations were successfully modeled by a meta-learning approach.
  • The model indicated that learning rate or inverse temperature was regulated by average reward rate.

Conclusions:

  • Rats progressively tune their reinforcement learning parameters during initial task learning.
  • Meta-learning provides a framework for understanding adaptive learning in uncertain environments.
  • Average reward rate is a key factor in regulating exploration-exploitation dynamics.