Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Operant Conditioning01:21

Operant Conditioning

1.5K
Operant conditioning, a key concept in behavioral psychology, involves using reinforcement and punishment to alter the likelihood of a behavior being repeated. B.F. introduced this type of conditioning. Skinner focused on voluntary behaviors and the consequences that follow them, influencing whether these behaviors will be strengthened or diminished.
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
1.5K
Reinforcement01:23

Reinforcement

154
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
154
Punishment01:27

Punishment

118
Negative reinforcement and punishment are often confused but serve distinct functions in behavior modification. Reinforcement, whether positive or negative, increases the likelihood of a desired behavior, while punishment decreases it.
Punishment can be positive or negative. Positive punishment involves adding an undesirable stimulus, such as scolding, to decrease a behavior. Negative punishment involves removing a desirable stimulus, such as taking away a favorite toy, to decrease behavior....
118
Generalization, Discrimination, and Extinction01:24

Generalization, Discrimination, and Extinction

337
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
337
Reinforcement Schedules01:24

Reinforcement Schedules

116
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
116
Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

57
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
57

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

MiR-21 promotes osteogenic transformation in ankylosing spondylitis fibroblasts and modulates bone metabolism in a murine arthritis model potentially involving the MAPK NF-κB pathway.

Journal of orthopaedic surgery and research·2026
Same author

A Machine Learning Driven Approach to Quantifying Coronary Artery Tortuosity.

JACC. Advances·2026
Same author

Is More Always Better With Digital Health Interventions? Shifting Engagement From Maximizing Use to Supporting Health.

Mayo Clinic proceedings. Digital health·2026
Same author

Elevated red cell distribution width as a prognostic indicator in critically ill patients with atrial fibrillation and chronic kidney disease.

BMC cardiovascular disorders·2026
Same author

Reproducible clinical archetypes in acute respiratory failure: a multi-cohort trajectory analysis.

Intensive care medicine·2026
Same author

Personalized modeling of stress and blood pressure reactivity using mobile health data.

Npj mental health research·2026

Related Experiment Video

Updated: May 5, 2026

Operant Procedures for Assessing Behavioral Flexibility in Rats
08:30

Operant Procedures for Assessing Behavioral Flexibility in Rats

Published on: February 15, 2015

20.9K

Rethinking Discount Regularization: New Interpretations, Unintended Consequences, and Solutions for Regularization in

Sarah Rathnam1, Sonali Parbhoo2, Siddharth Swaroop1

  • 1John A. Paulson School of Engineering and Applied Sciences, Harvard University, Cambridge, MA 02138 USA.

Journal of Machine Learning Research : JMLR
|May 8, 2025
PubMed
Summary

Discount regularization in reinforcement learning can lead to poor performance due to uneven data. This study introduces novel, state-action-specific methods to improve policy optimization by addressing these unintended consequences.

Keywords:
Markov decision processcertainty equivalencediscount factorregularizationreinforcement learning

More Related Videos

Measuring Delay Discounting in Humans Using an Adjusting Amount Task
07:47

Measuring Delay Discounting in Humans Using an Adjusting Amount Task

Published on: January 9, 2016

15.1K
New Variations for Strategy Set-shifting in the Rat
09:45

New Variations for Strategy Set-shifting in the Rat

Published on: January 23, 2017

7.7K

Related Experiment Videos

Last Updated: May 5, 2026

Operant Procedures for Assessing Behavioral Flexibility in Rats
08:30

Operant Procedures for Assessing Behavioral Flexibility in Rats

Published on: February 15, 2015

20.9K
Measuring Delay Discounting in Humans Using an Adjusting Amount Task
07:47

Measuring Delay Discounting in Humans Using an Adjusting Amount Task

Published on: January 9, 2016

15.1K
New Variations for Strategy Set-shifting in the Rat
09:45

New Variations for Strategy Set-shifting in the Rat

Published on: January 23, 2017

7.7K

Area of Science:

  • Artificial Intelligence
  • Machine Learning
  • Reinforcement Learning

Background:

  • Discount regularization is a common technique in reinforcement learning (RL) to prevent overfitting, especially with sparse or noisy data.
  • It is typically understood as reducing the importance of delayed rewards or future states in policy optimization.
  • Existing methods apply a global discount factor, which may not be optimal for all state-action pairs.

Purpose of the Study:

  • To reveal alternative interpretations of discount regularization that highlight its unintended consequences.
  • To motivate the development of novel regularization techniques that address these limitations.
  • To propose and validate state-action-specific regularization methods that generalize discount regularization.

Main Methods:

  • Proving two alternative theoretical views of discount regularization in both model-based and model-free RL.
  • Demonstrating how discount regularization acts as a prior favoring state-action pairs with more transition data in model-based RL.
  • Showing discount regularization's equivalence to a weighted average Bellman update in model-free RL.

Main Results:

  • Discount regularization can lead to poor performance when transition matrices are estimated from unevenly distributed data.
  • The proposed state-action-specific methods generalize discount regularization by adapting parameters locally.
  • Empirical validation across tabular and continuous state spaces demonstrates the effectiveness of the novel methods.

Conclusions:

  • Discount regularization has unintended consequences, particularly with varied data distributions across state-action pairs.
  • State-action-specific regularization offers a more robust and adaptable approach to policy optimization in RL.
  • The novel methods provide a remedy for the failures of standard discount regularization, improving performance in diverse RL settings.