DICCR:双门干预和混视觉语言导航的因果推理
Dongming Zhou1, Jinsheng Deng2, Zhengbin Pang1
1School of Computer Science, National University of Defense Technology, Deya Road, Changsha, 410003, Hunan, China.
概括
本研究引入了一种新的视觉语言导航 (VLN) 方法,使用因果推理来减少多模式偏差. DICCR模型通过解决视觉和文本指令之间的虚假相关性来提高导航性能.
科学领域:
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
- 计算机视觉 计算机视觉
背景情况:
- 视觉语言导航 (VLN) 代理必须将视觉场景与文本指令相关联,以便连续做出决策.
- 现有的方法往往忽略了多模式数据偏差和视觉和文本之间的虚假相关性.
- 因果关系为理解和缓解这些多模式关系提供了一个框架.
研究的目的:
- 提出一种新的视觉语言导航方法,DICCR,利用因果推理来解决多模式偏差.
- 通过交叉模式的因果推理来削弱视觉和文本模式之间的潜在虚假相关性.
- 通过减少对虚假关联的依赖来加强代理人的决策.
主要方法:
- 开发了使用交叉模式推理用于导航的混因子的因果图.
- 通过语义关系指导的前门和后门因果干预来减少偏见.
- 设计了一个联合的局部-全球因果关注模块和特征融合匹配算法 (FFM).
主要成果:
- DICCR模型在R2R,REVERIE和RxR基准数据集上显示了显著的改进.
- 在R2R数据集上,SPL增长了3.25%,SR增长了4.13%.
- 超越基线模型的性能,建立一个新的最先进的性能.
结论:
- 因果推理有效地减轻了VLN多模式数据中的虚假相关性和偏差.
- 拟议的DICCR模型为视觉语言导航任务提供了强大的和有效的解决方案.
- 这些发现强调了因果推理在开发更可靠的人工智能代理方面的重要性.
相关概念视频
Language and Cognition
321
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
321
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K


