通过层次强化学习学习发现视觉和语言导航的内在子目标
概括
我们介绍了DISH,这是一种用于视觉和语言导航的等级强化学习 (RL) 方法. DISH自主发现内在的子目标,改善概括,消除了对昂贵的轨迹注释的需求.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 计算机视觉 计算机视觉
背景情况:
- 视觉和语言导航代理通常使用模仿学习 (IL),这可能会导致过度拟合和由于教师偏见而导致糟糕的概括.
- 将IL与强化学习 (RL) 结合起来,可以提高概括性,但对于轨迹注释而言,成本很高.
研究的目的:
- 开发一种基于RL的新层次方法,DISH,它克服了概括的局限性,并且消除了视觉和语言导航中昂贵的标签注释的需要.
- 为了使代理人能够自主学习可概括的内在子目标,而没有预定义的语义约束.
主要方法:
- 一个层次化的RL框架,其中管理代理将导航任务分解为内在的子目标.
- 一个工作代理利用一个次目标驱动的注意力机制和一个历史意识的区分器 (HAD) 在一个缩小的状态空间中进行高效的动作预测.
- 采用自我监督,经理生成子目标来指导员工,从而避免需要标记的行动.
主要成果:
- 在房间到房间 (R2R) 数据集上,DISH方法在准确性和效率上显著优于基线方法.
- 历史意识歧视器 (HAD) 通过结合历史信息并提供内在奖励,有效地减轻了奖励稀疏性.
结论:
- 与现有的方法相比,DISH为视觉和语言导航提供了一个更具普遍性和成本效益的解决方案.
- 自主发现内在子目标是提高复杂导航任务中的代理性能的一个有希望的方向.
相关概念视频
Hierarchy of Motor Control
2.6K
The hierarchy of motor control refers to the different levels of organization and processing involved in controlling movement in the body. These levels range from higher cortical areas involved in planning and decision-making to lower spinal cord reflexes that respond automatically to external stimuli.
2.6K
Cognitive Learning
238
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
238


