Related Experiment Video
Updated: Jun 28, 2025

07:52
Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
8.7K
Boosting Reinforcement Learning via Hierarchical Game Playing With State Relay
IEEE Transactions on Neural Networks and Learning Systems
|April 22, 2024
Summary
This study introduces Hierarchical Reinforcement Learning based on Game playing with State Relay (HGR) to improve deep reinforcement learning (DRL) sample utilization and exploration. HGR enhances training efficiency in complex motion planning tasks.
Area of Science:
- Robotics and Artificial Intelligence
- Machine Learning
- Motion Planning
Background:
- Deep Reinforcement Learning (DRL) is widely applied in motion planning.
- Current DRL methods reset agent states after task completion, leading to low sample utilization and limited environment exploration.
- Initial training stages in DRL exhibit weak learning abilities, impacting efficiency in complex tasks.
Purpose of the Study:
- To propose a novel Hierarchical Reinforcement Learning (HRL) framework called Hierarchical learning based on Game playing with State Relay (HGR).
- To enhance sample utilization, expand environment exploration, and improve training efficiency in complex motion planning tasks.
Main Methods:
- Introduced an auxiliary penalty to regulate task difficulty.
- Developed a state relay mechanism to utilize intermediate agent states and expand low-level policy exploration.
- Implemented the HGR framework for training agents in complex environments.
Main Results:
- The state relay mechanism effectively utilizes intermediate states, expanding environment exploration.
- The proposed HGR algorithm improves sample utilization rates and mitigates sparse reward issues.
- Simulation tests on MazeBase and MuJoCo platforms demonstrate significant performance enhancements.
Conclusions:
- HGR framework effectively addresses limitations in current DRL for motion planning.
- The state relay mechanism and auxiliary penalty contribute to improved sample efficiency and exploration.
- HGR offers a significant benefit to the broader reinforcement learning (RL) field.
Related Concept Videos
Reinforcement
202
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
202
Reinforcement Schedules
144
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
144
Observational Learning
168
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
168
Primary and Secondary Reinforcers
247
In psychology, reinforcement is a key concept in behavior modification. B.F. Skinner demonstrated this with his experiments involving rats in what is known as a Skinner box. The rats learned to press a lever to receive food, a primary reinforcer that fulfilled their innate need for nourishment.
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
Effective reinforcers for humans vary depending on the individual and the context. Primary reinforcers, such as food, water, sleep, shelter, and pleasure, have inherent value and satisfy basic biological...
247
Associative Learning
350
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
350
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K

