高CPI-HRL:人类因果感知和推理驱动的等级强化学习
Bin Chen1, Zehong Cao2, Wolfgang Mayer2
1University of South Australia, Adelaide, SA, Australia; Xi'an Jiao Tong-Liverpool University, Jiangsu, China.
概括
本研究介绍了人类因果感知和推断驱动的等级强化学习 (HCPI-HRL),这是一种使用人类因果洞察力自动发现子目标的新方法. 高CPI-HRL提高了代理人培训效率和复杂环境中的适应能力.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 强化学习是一种强化学习.
背景情况:
- 层次增强学习 (HRL) 通常需要广泛的专家知识来定义子目标,限制其效率和适应性.
- 复杂和动态的环境给当前的HRL代理带来了挑战,因为训练和适应能力有限.
研究的目的:
- 开发一种新的方法,人类因果感知和推断驱动的等级强化学习 (HCPI-HRL),通过因果关系推断有效的子目标结构和关键对象.
- 加强HRL工作人员的探索方向,并促进跨任务的学习次目标结构的再利用.
- 克服对HRL专家知识的依赖,以提高培训效率和适应性.
主要方法:
- 拟议的HCPI-HRL具有双层架构:用于分配子目标的元控制器和用于执行子目标的基于近接政策优化 (PPO) 的下层.
- 利用人类引导的因果感知和推断来发现子目标结构并从动态环境状态中识别关键对象.
- 纳入稳定的因果关系,以指导内在奖励生成和代理探索.
主要成果:
- 在离散和连续控制环境中,HCPI-HRL与基准方法 (层次和相邻PPO) 相比表现优越.
- 该方法在培训效率,探索能力和学习政策的可转移性方面取得了显著改善.
- 实验验证实了以人为导向的因果模型在推断次目标关系和增强代理学习方面的有效性.
结论:
- 通过人类因果洞察,HCPI-HRL通过自动化子目标发现成功解决了传统HRL的局限性.
- 拟议的方法在动态环境中增强了代理商的能力,奖励稀少,为更具适应性和高效的HRL代理商铺平了道路.
- 这项研究强调了将因果推理与HRL集成为更复杂和更自主的人工智能系统的潜力.
更多相关视频
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
17.3K
07:43Author Spotlight: A Novel Setup to Conduct Naturalistic Laboratory Experiments with Real Human Actors in Scenarios
Published on: August 4, 2023
1.8K
相关概念视频
Cognitive Learning
136
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
136
Purposive Learning
96
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
96
Observational Learning
117
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
117
Real-World Application of Classical Conditioning
505
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
505
Inductive Reasoning
59.8K
Inductive reasoning is a form of logical thinking that uses related observations to arrive at a general conclusion. It is uncertain and operates in degrees to which the conclusions are credible. As such, inductive arguments can be weak or strong, rather than valid or invalid, and conclusions can be used to formulate testable, falsifiable hypotheses.
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
59.8K
Perception
424
Perception is a fundamental psychological process that enables individuals to organize, interpret, and consciously experience sensory information. This process is crucial for understanding and interacting with the world around us. It includes both bottom-up and top-down processing, each playing a distinct role in how we perceive our environment.
Bottom-up processing begins at the sensory level, where receptors detect external environmental stimuli. These could include the tactile sensation of...
Bottom-up processing begins at the sensory level, where receptors detect external environmental stimuli. These could include the tactile sensation of...
424
