Related Experiment Video
Updated: Aug 25, 2025

Measuring Delay Discounting in Humans Using an Adjusting Amount Task
Published on: January 9, 2016
Adaptive Discount Factor for Deep Reinforcement Learning in Continuing Tasks with Uncertainty
MyeongSeop Kim1,2, Jung-Su Kim1, Myoung-Su Choi2
1Research Center for Electrical and Information Technology, Department of Electrical and Information Engineering, Seoul National University of Science and Technology, Seoul 01811, Korea.
This study introduces an adaptive discount factor for reinforcement learning (RL) agents, improving consistent learning performance. The adaptive rule, based on the advantage function, enhances both on-policy and off-policy algorithms.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Reinforcement learning (RL) agents learn by maximizing discounted rewards.
- The discount factor critically impacts RL performance, especially under uncertainty.
- Constant discount factors can limit learning effectiveness in dynamic environments.
Purpose of the Study:
- To propose an adaptive rule for the discount factor in RL.
- To enhance consistent learning performance by adapting the discount factor based on the advantage function.
- To demonstrate the applicability of the adaptive rule in both on-policy and off-policy RL algorithms.
Main Methods:
- Developed an adaptive discount factor rule utilizing the advantage function.
- Integrated the adaptive rule into Proximal Policy Optimization (PPO) for on-policy learning.
- Integrated the adaptive rule into Soft Actor-Critic (SAC) for off-policy learning.
- Validated the approach on Tetris gameplay and robot manipulator motion planning.
Main Results:
- The proposed adaptive discount factor achieved comparable or superior performance to optimal constant discount factors.
- The adaptive method demonstrated effectiveness in both on-policy (PPO) and off-policy (SAC) scenarios.
- Consistent training performance was achieved across different tasks.
Conclusions:
- The adaptive discount factor rule effectively optimizes learning performance in RL.
- This method offers a robust alternative to manual discount factor tuning.
- The adaptive approach is applicable to various deep reinforcement learning problems.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
The Anchoring-and-Adjustment Heuristic
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:

