Related Experiment Video
Updated: Jun 7, 2025

Investigating Motor Skill Learning Processes with a Robotic Manipulandum
Published on: February 12, 2017
CoSD: Balancing behavioral consistency and diversity in unsupervised skill discovery.
Shuai Qing1, Yi Sun2, Kun Ding2
1School of Computer Science and Technology, Soochow University, Suzhou, 215006, China.
Constrained Skill Discovery (CoSD) balances skill diversity and behavioral consistency in hierarchical reinforcement learning. This method enhances skill stability and performance on complex tasks by preventing skills from focusing solely on exploration.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Robotics
Background:
- Hierarchical reinforcement learning (HRL) faces challenges with sparse rewards, making unsupervised skill discovery crucial.
- Existing skill discovery methods often prioritize diversity over internal skill consistency, leading to unstable skills.
- Overemphasis on exploration can result in skills with diffuse state visit distributions, lacking behavioral concentration.
Purpose of the Study:
- To introduce the Constrained Skill Discovery (CoSD) algorithm for balancing skill diversity and behavioral consistency.
- To address the limitations of previous methods that overemphasize skill diversity at the expense of stability.
- To improve the reliability and performance of learned skills in complex reinforcement learning environments.
Main Methods:
- CoSD integrates forward and reverse decomposition of mutual information to optimize skill learning.
- A maximum entropy policy is employed to maximize the information-theoretic objective.
- The algorithm enforces low internal state entropy for each skill, promoting behavioral consistency.
Main Results:
- Skills discovered by CoSD demonstrated more concentrated state visit distributions compared to other methods.
- CoSD achieved enhanced behavioral consistency and stability in learned skills.
- Skills exhibiting higher behavioral consistency led to superior performance in complex downstream tasks.
Conclusions:
- CoSD effectively balances skill diversity with behavioral consistency, overcoming limitations of prior unsupervised skill discovery approaches.
- The proposed method leads to more stable and reliable skills, crucial for practical applications in reinforcement learning.
- Enhanced skill consistency positively impacts performance in complex tasks, highlighting the importance of this constraint.
Related Concept Videos
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Observational Learning
Self-Discrepancy Theory
Distribution Reliability and Automation
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...

