Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Visual Agnosia01:12

Visual Agnosia

1.4K
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
1.4K
Visual System01:26

Visual System

2.1K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
2.1K
Components of Language01:24

Components of Language

866
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
866
Purposive Learning01:22

Purposive Learning

545
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
545
Indirect Motor Pathways01:22

Indirect Motor Pathways

3.7K
The indirect motor or extrapyramidal pathways originate in the brainstem, the lower portion of the brain that connects it to the spinal cord. They consist of several distinct tracts, each with specialized functions. The four main tracts of the indirect motor pathways are the vestibulospinal tract, the reticulospinal tract, the tectospinal tract, and the rubrospinal tract.
The vestibulospinal tract originates in the vestibular nuclei of the brainstem. The vestibular system detects changes in...
3.7K
Fluid Mosaic Model01:19

Fluid Mosaic Model

18.5K
Scientists identified the plasma membrane in the 1890s and its principal chemical components (lipids and proteins) by 1915. The model for plasma membrane structure, proposed in 1935 by Hugh Davson and James Danielli, was the first model to be widely accepted in the scientific community. The model was based on the plasma membrane's "railroad track" appearance in early electron micrographs. Davson and Danielli theorized that the plasma membrane's structure resembled a sandwich...
18.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

SR-VLN: Implicit Spatial Reasoning Vision-and-Language Navigation.

Sensors (Basel, Switzerland)·2026
Same author

[The effect of cyclophosphamide on cytokines in patients with primary Sjögren's syndrome-associated interstitial lung disease].

Zhonghua jie he he hu xi za zhi = Zhonghua jiehe he huxi zazhi = Chinese journal of tuberculosis and respiratory diseases·2011
Same author

Seroprevalence of Toxoplasma gondii infection in slaughtered pigs and cattle in Liaoning Province, northeastern China.

The Journal of parasitology·2011
Same author

Evaluating the effectiveness of a schools-based programme to promote exercise self-efficacy in children and young people with risk factors for obesity: steps to active kids (STAK).

BMC public health·2011
Same author

Survival advantage of normal weight in peritoneal dialysis patients.

Renal failure·2011
Same author

[Application value of combining brain natriuretic peptide, creatine phosphokinase and echocardiogram in the evaluation of polymyositis-related chronic heart failure].

Sichuan da xue xue bao. Yi xue ban = Journal of Sichuan University. Medical science edition·2011

Related Experiment Video

Updated: Feb 28, 2026

Practical Methodology of Cognitive Tasks Within a Navigational Assessment
05:19

Practical Methodology of Cognitive Tasks Within a Navigational Assessment

Published on: June 1, 2015

14.1K

CA-VLN: Collaborative Agents in MLLM-Powered Visual-Language Navigation.

Ruolin Zhu1, Shaobin Li1, Zixing Zhu1

  • 1School of Information and Communication Engineering, Communication University of China, Beijing 100024, China.

Sensors (Basel, Switzerland)
|February 27, 2026
PubMed
Summary

This study introduces Collaborative Agents in Visual-Language Navigation (CA-VLN), a novel framework using multimodal large language models to improve robot navigation in unseen environments. CA-VLN enhances generalization and success rates through collaborative agents for semantic understanding and memory-based planning.

Keywords:
episodic memoryhierarchical historymultimodal feature fusionworld knowledge

More Related Videos

Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions
07:09

Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions

Published on: May 2, 2019

6.5K
A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
06:28

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants

Published on: August 26, 2018

6.3K

Related Experiment Videos

Last Updated: Feb 28, 2026

Practical Methodology of Cognitive Tasks Within a Navigational Assessment
05:19

Practical Methodology of Cognitive Tasks Within a Navigational Assessment

Published on: June 1, 2015

14.1K
Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions
07:09

Integrating Visual Psychophysical Assays within a Y-Maze to Isolate the Role that Visual Features Play in Navigational Decisions

Published on: May 2, 2019

6.5K
A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
06:28

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants

Published on: August 26, 2018

6.3K

Area of Science:

  • Robotics
  • Artificial Intelligence
  • Computer Vision

Background:

  • Generalizing robot navigation to new, unseen environments is a significant challenge in Vision-Language Navigation (VLN).
  • Existing methods often struggle with long-horizon planning and integrating commonsense reasoning for robust navigation.

Purpose of the Study:

  • To develop a novel framework that enhances generalization capabilities in Vision-Language Navigation.
  • To leverage world knowledge from Multimodal Large Language Models for improved navigation performance.
  • To enable robots to navigate successfully in previously unobserved environments.

Main Methods:

  • Introduction of Collaborative Agents in Visual-Language Navigation (CA-VLN), a dual-agent architecture.
  • A Knowledge Agent integrates semantic context and commonsense reasoning into action prediction.
  • A Hierarchical History Agent builds detailed episodic memory for long-horizon planning.

Main Results:

  • CA-VLN achieves state-of-the-art performance on established benchmarks (R2R, REVERIE, SOON).
  • Significant improvements in generalization to unseen environments were observed.
  • Enhanced navigation success rates in novel and complex settings.

Conclusions:

  • The proposed CA-VLN framework effectively utilizes multimodal large language models for superior navigation.
  • The dual-agent architecture enables a dynamic interplay between semantic understanding and episodic memory.
  • CA-VLN represents a significant advancement in addressing the generalization challenge in Vision-Language Navigation.