CoSD:在无监督技能发现中平衡行为一致性和多样性
Shuai Qing1, Yi Sun2, Kun Ding2
1School of Computer Science and Technology, Soochow University, Suzhou, 215006, China.
概括
约束技能发现 (CoSD) 在层次强化学习中平衡技能多样性和行为一致性. 这种方法提高了技能稳定性和复杂任务的性能,通过防止技能只专注于探索.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 机器人技术 机器人技术 机器人技术
背景情况:
- 层次强化学习 (HRL) 面临着奖励稀少的挑战,使得无监督的技能发现至关重要.
- 现有的技能发现方法往往将多样性优先于内部技能的一致性,导致不稳定的技能.
- 对探索的过度强调可能会导致状态访问分布分散的技能,缺乏行为集中度.
研究的目的:
- 引入受限技能发现 (CoSD) 算法,以平衡技能多样性和行为一致性.
- 解决以前方法的局限性,这些方法过度强调技能多样性而牺牲了稳定性.
- 在复杂的强化学习学习环境中提高学习技能的可靠性和性能.
主要方法:
- CoSD集成了相互信息的前向和反向分解,以优化技能学习.
- 使用最大率政策来最大限度地实现信息理论的目标.
- 该算法强制执行每个技能的低内部状态,促进行为一致性.
主要成果:
- 与其他方法相比,CoSD发现的技能显示了比其他方法更集中的国家访问分布.
- CoSD实现了增强的行为一致性和学习技能的稳定性.
- 表现出更高行为一致性的技能导致在复杂的下游任务中表现出色.
结论:
- CoSD有效地平衡了技能多样性与行为一致性,克服了以前无监督技能发现方法的局限性.
- 拟议的方法导致更稳定和可靠的技能,对于强化学习的实际应用至关重要.
- 提高技能一致性对复杂任务的表现产生了积极的影响,突出了这一约束的重要性.
相关概念视频
Generalization, Discrimination, and Extinction
445
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
445
Observational Learning
135
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
135
Self-Discrepancy Theory
18.3K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.
18.3K
Distribution Reliability and Automation
105
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
105
Cognitive Learning
220
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
220
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K


