Related Experiment Video
Updated: Jan 8, 2026

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
Towards more effective skill discovery in reinforcement learning by incorporating state reachability
Yang Liu1, Jingchen Li2, Huarui Wu2
1College of Optical Science and Engineering, Zhejiang University, Zhejiang Province, Hangzhou, 310058, China.
None:
Skill discovery in reinforcement learning seeks to autonomously learn a diverse repertoire of behaviors, enabling efficient adaptation to downstream tasks. While existing methods primarily maximize mutual information between skills and states to ensure diversity, they often fail to guarantee state reachability, leading to skill-unreachable regions that hinder adaptation in complex environments. To address this, we propose Skill Discovery with State Reachability (SDSR), a novel framework that explicitly integrates reachability into skill learning. SDSR enhances traditional mutual information-based methods by introducing a skill-conditioned inverse dynamics model, which learns the necessary state-action transitions to expand the agent's accessible state space, and a meta-policy optimization mechanism, which jointly optimizes skill diversity and reachability to ensure comprehensive state coverage. We implement SDSR through two complementary approaches: threshold-based selection, which leverages high-probability state-action pairs from experience for efficient skill learning in low-dimensional environments, and joint training, which optimizes reachability alongside reinforcement learning objectives, making it well-suited for high-dimensional environments. Extensive experiments in both 2D and robotic environments demonstrate that SDSR significantly improves skill diversity, enhances exploration efficiency, and accelerates adaptation to downstream tasks. By expanding the agent's accessible state space while maintaining structured skill diversity, SDSR provides a robust and generalizable foundation for reinforcement learning in complex decision-making domains.
Related Concept Videos
Observational Learning
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
State Space Representation
Consider an RLC circuit, a...
Role of Shaping in Operant Conditioning
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...

