Related Experiment Video
Updated: Jul 17, 2025

A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence
Sarah Rathnam1, Sonali Parbhoo2, Weiwei Pan1
1Harvard University, School of Engineering and Applied Sciences, Cambridge, MA USA.
Discount regularization in Markov Decision Processes (MDPs) can lead to poor policy estimation with uneven data. This study reveals discount regularization acts like a state-action prior, proposing a new method for state-action-specific regularization parameters.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Reinforcement Learning
Background:
- Discount regularization is commonly used in reinforcement learning to simplify policy estimation from sparse or noisy data.
- It is generally understood to de-emphasize delayed effects in Markov Decision Processes (MDPs).
Purpose of the Study:
- To reveal an alternative perspective on discount regularization, highlighting its unintended consequences.
- To demonstrate that discount regularization is equivalent to a prior on the transition matrix.
- To propose a novel method for setting state-action-specific regularization parameters.
Main Methods:
- Developed an equivalence theorem showing discount regularization is equivalent to a prior on the transition matrix.
- Derived an explicit formula for state-action-specific regularization parameters.
- Empirically evaluated the proposed method against standard discount regularization using simulations and a medical cancer simulator.
Main Results:
- Demonstrated that discount regularization functions as a prior with stronger regularization on state-action pairs with more data.
- Showcased that this leads to suboptimal performance when transition matrices are estimated from uneven datasets.
- The proposed state-action-specific method significantly improves policy estimation accuracy.
Conclusions:
- Discount regularization has unintended consequences, acting as a data-dependent prior.
- A state-action-specific regularization approach remedies the shortcomings of global discount regularization.
- The proposed method offers improved policy estimation in reinforcement learning, especially with uneven data.
More Related Videos
05:22Dissociation of the Confounding Influences of Expectancy and Integrative Difficulty Residing in Anomalous Sentences in Event-related Potential Studies
Published on: May 9, 2019
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Propagation of Uncertainty from Random Error
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Propagation of Uncertainty from Systematic Error
Constraints and Statical Determinacy
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Hindsight Biases