以历史为导向的快速生成用于视觉和语言导航
IEEE transactions on cybernetics
|October 2, 2025
概括
本研究介绍了视觉和语言导航 (VLN) 的历史引导提示生成 (HGPG) 框架. 该方法适应地挖掘历史数据,以改善新环境中的代理人感知.
科学领域:
- 化身的人工智能
- 计算机视觉 计算机视觉 计算机视觉
- 自然语言处理自然语言处理.
背景情况:
- 视觉和语言导航 (VLN) 依赖于历史观察以获得上下文知识.
- 目前的VLN方法未能明确地将历史背景与当前环境联系起来.
- 在现有的方法中,对环境特定线索的自适应性学习往往被忽视.
研究的目的:
- 通过自适应地挖掘相关的历史信息来增强视觉和语言导航中的代理人感知.
- 为VLN提出一个新的历史导向提示生成 (HGPG) 框架.
- 通过跨任务共享已学到的表示来改进对未知的环境的概括.
主要方法:
- 开发了一个基于的历史获取模块,以评估历史信息的必要性.
- 实现了一个提示生成模块,该模块使用学习的令牌库将历史上下文转换为紧的提示向量.
- 在各种导航任务中使用共享的代币库,以捕捉共同特征并增强概括性.
主要成果:
- 在四个主流的VLN基准 (R2R,REVERIE,SOON,R2R-CE) 上,HGPG框架表现出显著的有效性.
- 拟议的方法通过利用历史数据,成功地提高了代理人对当前环境的感知.
- 分享代币库改善了对以前未见的环境的概括能力.
结论:
- 以历史为导向的提示生成框架为改善视觉和语言导航提供了一个有希望的方法.
- 适应性挖掘历史信息对于稳健和通用导航至关重要.
- 在动态环境中,HGPG方法为代理人提供了一种更有效和更有效的方式来利用过去的经验.
相关概念视频
Vision
59.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.4K
Visual Agnosia
974
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
974
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K
Visual System
1.7K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.7K


