Related Experiment Video
Updated: Sep 23, 2025

09:01
The Double-H Maze: A Robust Behavioral Test for Learning and Memory in Rodents
Published on: July 8, 2015
12.7K
Hierarchical Reinforcement Learning, Sequential Behavior, and the Dorsal Frontostriatal System
Miriam Janssen1, Christopher LeWarne1, Diana Burk1
1National Institute of Mental Health, Bethesda, MD.
Journal of Cognitive Neuroscience
|May 17, 2022
Summary
Biological agents learn complex tasks by breaking them down. This review suggests human motor sequences, like those in visuomotor tasks, can model options in hierarchical reinforcement learning (HRL) and their brain mechanisms.
Area of Science:
- Cognitive Neuroscience
- Computational Neuroscience
- Motor Control
Background:
- Biological agents require adaptive behavior in dynamic environments, necessitating learning and action at multiple hierarchical levels.
- Hierarchical reinforcement learning (HRL) models this by creating temporally extended actions ('options') from sequences.
- Key questions in HRL concern option formation and neural realization.
Purpose of the Study:
- To explore how human motor sequence learning literature can inform understanding of option formation in HRL.
- To investigate the neural mechanisms underlying HRL through the lens of motor control.
- To bridge insights from motor sequence learning and reinforcement learning.
Main Methods:
- Review of existing human motor sequence literature, focusing on visuomotor tasks (e.g., discrete sequence production, M x N task).
- Analysis of how hierarchical learning and behavior are represented in sequential action tasks.
- Examination of the potential role of dorsal cortical-subcortical circuitry in supporting HRL.
Main Results:
- Motor chunks within learned sequences can be conceptualized as HRL options.
- Visuomotor sequence learning tasks provide a framework for studying hierarchical behavior.
- The dorsal cortical-subcortical pathway is implicated in supporting such hierarchical processes.
Conclusions:
- Human motor sequence learning offers valuable insights into the formation and neural basis of options in HRL.
- Integrating motor sequence literature with RL can advance experimental designs in both fields.
- This interdisciplinary approach can elucidate how complex behaviors are learned and executed hierarchically.
Related Concept Videos
Reinforcement Schedules
244
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
244
Real-World Application of Classical Conditioning
766
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
766
Associative Learning
612
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
612
Timing and Consequences on Behavior
161
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
161
Generalization, Discrimination, and Extinction
836
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
836
Law of Effect
1.7K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
1.7K

