Related Experiment Video
Updated: Jan 10, 2026

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
Hit the spot: Reachability guided subgoal generation for hierarchical reinforcement learning in stochastic
Bo Chen1, Quan Yuan1, Guiyang Luo1
1Beijing University of Posts and Telecommunications china, China.
None:
Hierarchical reinforcement learning (HRL) with subgoals offers an effective approach to tackling tasks characterized by sparse rewards. HRL enables agents to plan across multiple levels where high levels generate subgoals and low levels pursue to reach them. However, execution ability of low level is unstable during training, particularly in stochastic environments, resulting in inaccurate expectations by high level regarding the reachability of subgoals. To address this problem, we propose a reachability guided subgoal generation method based on the representation of the low level's execution ability. Firstly, we theoretically prove that the shortest transition distance between the arrival state and the subgoal follows a normal distribution in both deterministic and stochastic environments, and derive its mean and variance. The high level utilizes the distance model to represent the abilities of the low level, which enhances its capacity to generate subgoals. Therefore, the distance model is incorporated into the value prediction network architecture of the high level, ultimately guiding the policy network to generate well-fitting subgoals for the low level in stochastic environments. Experimental results indicate that our approach achieves higher completion and faster convergence rate in stochastic environments.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Statically Indeterminate Problem Solving
Reinforcement Schedules
Once a behavior is learned,...
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Hydraulic Jump: Problem Solving

