Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reinforcement Schedules01:24

Reinforcement Schedules

230
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
230
Timing and Consequences on Behavior01:08

Timing and Consequences on Behavior

144
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective. 
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
144
Avoidance Learning and Learned Helplessness01:14

Avoidance Learning and Learned Helplessness

1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K
Decision Making: P-value Method01:09

Decision Making: P-value Method

5.6K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim  is also stated. These statements can act as null and alternative hypotheses:  a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.6K
The Anchoring-and-Adjustment Heuristic01:25

The Anchoring-and-Adjustment Heuristic

7.4K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.4K
Reinforcement01:23

Reinforcement

313
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
313

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Path Planning for Multi-Arm Manipulators Using Deep Reinforcement Learning: Soft Actor-Critic with Hindsight Experience Replay.

Sensors (Basel, Switzerland)·2020
See all related articles

Related Experiment Video

Updated: Aug 25, 2025

Measuring Delay Discounting in Humans Using an Adjusting Amount Task
07:47

Measuring Delay Discounting in Humans Using an Adjusting Amount Task

Published on: January 9, 2016

15.5K

Adaptive Discount Factor for Deep Reinforcement Learning in Continuing Tasks with Uncertainty.

MyeongSeop Kim1,2, Jung-Su Kim1, Myoung-Su Choi2

  • 1Research Center for Electrical and Information Technology, Department of Electrical and Information Engineering, Seoul National University of Science and Technology, Seoul 01811, Korea.

Sensors (Basel, Switzerland)
|October 14, 2022
PubMed
Summary

This study introduces an adaptive discount factor for reinforcement learning (RL) agents, improving consistent learning performance. The adaptive rule, based on the advantage function, enhances both on-policy and off-policy algorithms.

Keywords:
Tetrisdiscount factorpath planningreinforcement learninguncertainty

More Related Videos

Errors as a Means of Reducing Impulsive Food Choice
07:07

Errors as a Means of Reducing Impulsive Food Choice

Published on: June 5, 2016

8.7K
Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats
09:12

Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats

Published on: March 17, 2019

9.5K

Related Experiment Videos

Last Updated: Aug 25, 2025

Measuring Delay Discounting in Humans Using an Adjusting Amount Task
07:47

Measuring Delay Discounting in Humans Using an Adjusting Amount Task

Published on: January 9, 2016

15.5K
Errors as a Means of Reducing Impulsive Food Choice
07:07

Errors as a Means of Reducing Impulsive Food Choice

Published on: June 5, 2016

8.7K
Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats
09:12

Three Laboratory Procedures for Assessing Different Manifestations of Impulsivity in Rats

Published on: March 17, 2019

9.5K

Area of Science:

  • Artificial Intelligence
  • Machine Learning
  • Robotics

Background:

  • Reinforcement learning (RL) agents learn by maximizing discounted rewards.
  • The discount factor critically impacts RL performance, especially under uncertainty.
  • Constant discount factors can limit learning effectiveness in dynamic environments.

Purpose of the Study:

  • To propose an adaptive rule for the discount factor in RL.
  • To enhance consistent learning performance by adapting the discount factor based on the advantage function.
  • To demonstrate the applicability of the adaptive rule in both on-policy and off-policy RL algorithms.

Main Methods:

  • Developed an adaptive discount factor rule utilizing the advantage function.
  • Integrated the adaptive rule into Proximal Policy Optimization (PPO) for on-policy learning.
  • Integrated the adaptive rule into Soft Actor-Critic (SAC) for off-policy learning.
  • Validated the approach on Tetris gameplay and robot manipulator motion planning.

Main Results:

  • The proposed adaptive discount factor achieved comparable or superior performance to optimal constant discount factors.
  • The adaptive method demonstrated effectiveness in both on-policy (PPO) and off-policy (SAC) scenarios.
  • Consistent training performance was achieved across different tasks.

Conclusions:

  • The adaptive discount factor rule effectively optimizes learning performance in RL.
  • This method offers a robust alternative to manual discount factor tuning.
  • The adaptive approach is applicable to various deep reinforcement learning problems.