Related Experiment Video
Updated: Sep 17, 2025

11:20
Recording Single Neurons' Action Potentials from Freely Moving Pigeons Across Three Stages of Learning
Published on: June 2, 2014
12.1K
Ephemeral reward task: Why is it so difficult for pigeons to learn it?
Daniel Peng1, Zohaib Iqbal1, Thomas R Zentall1
1Department of Psychology, University of Kentucky.
Summary
Pigeons struggled with the Ephemeral Reward Task until reward probabilities were reduced. Lowering reward probability to 50% enabled pigeons to learn optimal choice strategies.
Area of Science:
- Animal Cognition
- Behavioral Neuroscience
- Comparative Psychology
Background:
- The Ephemeral Reward Task involves choosing between stimuli A and B for rewards.
- Wrasse, parrots, and primates learn this task, but pigeons and rats do not.
- Previous research suggests pigeons' difficulty may stem from stimulus-outcome similarity.
Purpose of the Study:
- To investigate why pigeons fail to learn the Ephemeral Reward Task.
- To test the hypothesis that stimulus-outcome similarity hinders learning.
- To explore the role of reward magnitude and probability in pigeon learning.
Main Methods:
- Experiment 1: Modified stimuli (C) after choice B to test outcome similarity hypotheses (groups AC, BC, BB).
- Experiment 2: Reduced reward probability for choices A and B to 50% for pigeons.
- Behavioral observation and analysis of choice patterns to assess learning.
Main Results:
- Pigeons in Experiment 1 failed to learn optimal strategies across all modified conditions (AC, BC, BB).
- Pigeons in Experiment 2 successfully learned to choose optimally when reward probability was 50%.
- The difference in perceived value between one and two rewards might be less impactful than the difference between 0.5 and one reward.
Conclusions:
- Stimulus-outcome similarity does not appear to be the primary reason for pigeons' difficulty in the standard Ephemeral Reward Task.
- Reward probability significantly influences pigeons' ability to learn optimal strategies in this task.
- The relative value of rewards, particularly the difference between fractional and certain rewards, is a critical factor in learning.
Related Concept Videos
Timing and Consequences on Behavior
157
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
157
Instinctive Drift
330
Instinctive drift refers to the tendency of animals to revert to their innate behaviors despite repeated reinforcement. Breland and Breland demonstrated this concept in an experiment with a raccoon. The raccoon was trained to pick up two coins and place them in a container in exchange for food. Initially, the raccoon learned to associate the coins with food, making them a conditioned stimulus or a substitute for food. However, over time, the raccoon became less willing to put the coins into the...
330

