Related Experiment Video
Updated: Jun 4, 2026

The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents
Published on: July 8, 2015
Ventral striatum and orbitofrontal cortex are both required for model-based, but not model-free, reinforcement
Michael A McDannald1, Federica Lucantonio, Kathryn A Burke
1Department of Anatomy and Neurobiology, University of Maryland School of Medicine, Baltimore, Maryland 21201, USA. mmcda001@umaryland.edu
Learning can occur from unexpected reward identity, not just value changes. This suggests reinforcement learning models need to incorporate environmental context and identity-based prediction errors, involving the ventral striatum and orbitofrontal cortex.
Area of Science:
- Neuroscience
- Behavioral Science
- Computational Neuroscience
Background:
- Learning is often attributed to reward value prediction errors.
- However, learning can also occur when reward identity changes, even if value remains constant.
- This identity-based learning suggests a need for internal environmental models, challenging traditional model-free reinforcement learning.
Purpose of the Study:
- To investigate the neural mechanisms underlying value-based versus identity-based prediction error learning.
- To differentiate the roles of brain regions in learning from changes in reward value versus identity.
- To inform and potentially modify existing reinforcement learning models.
Main Methods:
- Utilized unblocking procedures in rats to assess learning.
- Trained rats to associate visual cues with specific food quantities and identities.
- Introduced novel auditory cues paired with changes in reward quantity or identity, followed by probe tests.
Main Results:
- The ventral striatum was essential for learning from changes in reward value, consistent with temporal difference reinforcement learning (TDRL) models.
- Both the ventral striatum and orbitofrontal cortex were required for learning from changes in reward identity.
- These findings indicate that identity-based learning relies on more than just value prediction errors.
Conclusions:
- Existing TDRL models need to be updated to incorporate model-based representations of outcome features for identity-based learning.
- The orbitofrontal cortex plays a crucial role in identity-based learning, requiring further delineation within these models.
- This research highlights the complexity of learning mechanisms beyond simple reward valuation.
Related Concept Videos
Observational Learning
Association Areas of the Cortex
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
Associative Learning
Classical conditioning, also known...
Operant Conditioning
Reinforcement in operant conditioning can be positive or negative, both of which serve to increase the likelihood of a behavior. Positive...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Steps in the Modeling Process
Attention is the first necessary component for observational learning. It involves focusing on what the model is doing and saying. For example, if you decide to take a drawing class to enhance your skills, you need to pay close attention to the instructor's words and hand movements. The characteristics of the model significantly...

