Related Experiment Video
Updated: Mar 17, 2026

Measuring Delay Discounting in Humans Using an Adjusting Amount Task
Published on: January 9, 2016
When and Why Hyperbolic Discounting Matters for Reinforcement Learning Interventions
Ian M Moore1, Eura Nofshin1, Siddharth Swaroop1
1Department of Computer Science, Harvard University, USA.
Abstract:
In settings where an AI agent nudges a human agent toward a goal, the quality of the AI's policy depends on how well it models the human. Despite behavioral evidence that humans hyperbolically discount future rewards, the RL community continues to model humans as Markov Decision Processes (MDPs) with exponential discounting. This is because planning is difficult with non-exponential discounts. In this work, we investigate whether the performance benefits of modeling humans as hyperbolic discounters outweigh the computational costs. We focus on AI interventions that change the human's discounting (i.e. decreases the human's "nearsightedness" to help them toward distant goals). We derive a fixed exponential discount factor that can approximate hyperbolic discounting, and prove that this approximation guarantees the AI will never miss a necessary intervention. We also prove that our approximation causes fewer false positives (unnecessary interventions) than the mean hazard rate, another well-known method for approximating hyperbolic MDPs as exponential ones. Surprisingly, our experiments demonstrate that exponential approximations outperform hyperbolic ones in online learning, even when the ground-truth human MDP is hyperbolically discounted.
Related Concept Videos
Operant Conditioning Intervention
In operant conditioning, behaviors that are...
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Primary and Secondary Reinforcers
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...

