Related Experiment Video
Updated: Feb 28, 2026

Practical Methodology of Cognitive Tasks Within a Navigational Assessment
Published on: June 1, 2015
CA-VLN: Collaborative Agents in MLLM-Powered Visual-Language Navigation
Ruolin Zhu1, Shaobin Li1, Zixing Zhu1
1School of Information and Communication Engineering, Communication University of China, Beijing 100024, China.
Abstract:
Generalization to unseen environments remains a fundamental challenge in Vision-Language Navigation. To tackle this issue, we propose a novel framework that leverages world knowledge embedded within Multimodal Large Language Models. We introduce Collaborative Agents in Visual-Language Navigation (CA-VLN), a framework based on a dual-agent architecture. This architecture comprises a Knowledge Agent, which infuses the action prediction process with semantic context and commonsense reasoning, and a Hierarchical History Agent, which constructs a detailed episodic memory to enable long-horizon planning. The collaboration between these agents facilitates a dynamic interplay between high-level semantic understanding and grounded episodic experience. Extensive experiments on the R2R, REVERIE and SOON datasets demonstrate that our model achieves state-of-the-art performance, significantly improving generalization and navigation success in previously unobserved environments.
More Related Videos
07:09Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions
Published on: May 2, 2019
06:28A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
Related Concept Videos
Visual Agnosia
Visual System
Once through the pupil, the light passes through the lens, a...
Components of Language
Purposive Learning
Indirect Motor Pathways
The vestibulospinal tract originates in the vestibular nuclei of the brainstem. The vestibular system detects changes in...
Fluid Mosaic Model