DICCR: Double-gated intervention and confounder causal reasoning for vision-language navigation

Dongming Zhou1, Jinsheng Deng2, Zhengbin Pang1

  • 1School of Computer Science, National University of Defense Technology, Deya Road, Changsha, 410003, Hunan, China.

Summary

This study introduces a novel approach for vision-language navigation (VLN) using causal reasoning to reduce multi-modal bias. The DICCR model improves navigation performance by addressing spurious correlations between vision and text instructions.

Related Concept Videos