Related Experiment Video
Updated: Jul 1, 2026

Measuring Delay Discounting in Humans Using an Adjusting Amount Task
Published on: January 9, 2016
Delayed reward information is underweighted in reinforcement learning with dispersed feedback
Miruna Cotet1,2, David Poensgen3, Ian Krajbich1,4,5
1Department of Psychology, The Ohio State University, Columbus, Ohio, United States of America.
People learning from choices with immediate and delayed reward information overweighted immediate feedback, leading to worse outcomes. This bias persisted even when learning from others, indicating a preference for immediate information over actual reward timing.
Area of Science:
- Cognitive psychology
- Neuroeconomics
- Decision-making
Background:
- Adaptive behavior relies on learning from action outcomes.
- Traditional learning models often assume single outcomes per action.
- Real-world scenarios involve actions with multiple, temporally distributed consequences.
Purpose of the Study:
- To investigate how individuals learn when faced with both immediate and delayed reward information.
- To determine if reward timing influences information weighting, irrespective of actual reward delivery.
- To explore the persistence and contributing factors of biased information processing.
Main Methods:
- Utilized behavioral experiments to track learning and decision-making.
- Employed eye-tracking to analyze visual attention towards immediate versus delayed feedback.
- Designed tasks where all rewards were delivered post-experiment to isolate information processing biases.
Main Results:
- Subjects consistently overweighted immediate reward information, despite equivalent reward delivery.
- This bias intensified throughout the experiment and was observed even when learning socially.
- Eye-tracking data showed inconclusive evidence linking gaze patterns to the observed behavioral bias.
Conclusions:
- Individuals exhibit a bias towards prioritizing immediate reward *information*, not just immediate rewards.
- This preference for immediacy in information processing is a cognitive error, distinct from temporal discounting.
- The findings suggest a fundamental mechanism in human learning that can lead to suboptimal decision-making.
More Related Videos
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Reinforcement Schedules
Once a behavior is learned,...
Timing and Consequences on Behavior
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant factor...
Instinctive Drift

