Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
Vision01:24

Vision

Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
Visual Agnosia01:12

Visual Agnosia

Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round end"...
Visual System01:26

Visual System

Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Parallel Processing01:20

Parallel Processing

The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
Motor and Sensory Areas of the Cortex01:14

Motor and Sensory Areas of the Cortex

The cerebral cortex, the brain's outermost layer, is pivotal in processing complex cognitive tasks, emotions, and various sensory inputs and executing voluntary motor activities. This intricate structure is divided into three primary functional areas: the motor areas, sensory areas, and association areas.
Motor Areas
The motor areas located in the frontal lobe are central to controlling voluntary movements. This region is further subdivided into the primary motor cortex and the premotor cortex.

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

CA-VLN: Collaborative Agents in MLLM-Powered Visual-Language Navigation.

Sensors (Basel, Switzerland)·2026
Same author

[The effect of cyclophosphamide on cytokines in patients with primary Sjögren's syndrome-associated interstitial lung disease].

Zhonghua jie he he hu xi za zhi = Zhonghua jiehe he huxi zazhi = Chinese journal of tuberculosis and respiratory diseases·2011
Same author

Seroprevalence of Toxoplasma gondii infection in slaughtered pigs and cattle in Liaoning Province, northeastern China.

The Journal of parasitology·2011
Same author

Evaluating the effectiveness of a schools-based programme to promote exercise self-efficacy in children and young people with risk factors for obesity: steps to active kids (STAK).

BMC public health·2011
Same author

Survival advantage of normal weight in peritoneal dialysis patients.

Renal failure·2011
Same author

[Application value of combining brain natriuretic peptide, creatine phosphokinase and echocardiogram in the evaluation of polymyositis-related chronic heart failure].

Sichuan da xue xue bao. Yi xue ban = Journal of Sichuan University. Medical science edition·2011

Related Experiment Video

Updated: Jun 27, 2026

Assessing Human Spatial Navigation in a Virtual Space and its Sensitivity to Exercise
06:17

Assessing Human Spatial Navigation in a Virtual Space and its Sensitivity to Exercise

Published on: January 26, 2024

SR-VLN: Implicit Spatial Reasoning Vision-and-Language Navigation.

Ruolin Zhu1, Shaobin Li1, Min Yang2

  • 1School of Information and Communication Engineering, Communication University of China, Beijing 100024, China.

Sensors (Basel, Switzerland)
|June 26, 2026
PubMed
Summary

Spatial Reasoning Vision-and-Language Navigation (SR-VLN) enhances efficiency by using implicit spatial tokens instead of explicit reasoning. This approach significantly speeds up navigation and reduces token usage for real-time performance.

Keywords:
implicit reasoningmultimodal fusionperceptual compressionpyramidal hierarchical historyspatial reasoning

More Related Videos

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
06:28

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants

Published on: August 26, 2018

Related Experiment Videos

Last Updated: Jun 27, 2026

Assessing Human Spatial Navigation in a Virtual Space and its Sensitivity to Exercise
06:17

Assessing Human Spatial Navigation in a Virtual Space and its Sensitivity to Exercise

Published on: January 26, 2024

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
06:28

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants

Published on: August 26, 2018

Area of Science:

  • Artificial Intelligence
  • Robotics
  • Computer Vision

Background:

  • Traditional Vision-and-Language Navigation (VLN) relies on explicit reasoning, limiting efficiency and scalability in complex environments.
  • Multimodal Large Language Models (MLLMs) face latency issues due to verbose textual outputs during decision-making.

Purpose of the Study:

  • To introduce Spatial Reasoning Vision-and-Language Navigation (SR-VLN), a novel framework for efficient and scalable VLN.
  • To shift from explicit chain-of-thought (CoT) reasoning to an implicit spatial representation space.

Main Methods:

  • Developed a pyramidal hierarchical history framework with perceptual compression to create multi-scale trajectory representations.
  • Introduced compact, learnable spatial tokens (S-Tokens) for agile inference in the latent feature space.
  • Employed a hybrid training strategy combining sparse reward supervision and Proximal Policy Optimization (PPO) reinforcement learning.

Main Results:

  • SR-VLN achieved state-of-the-art navigation performance across R2R, REVERIE, and SOON datasets.
  • Reduced token consumption by 68% and achieved a 4.1x inference speedup compared to explicit reasoning baselines.
  • Attained a 76.02% success rate and 73.80% SPL on the R2R unseen split.

Conclusions:

  • SR-VLN offers a significant improvement in efficiency and scalability for long-range navigation tasks.
  • The implicit spatial representation approach facilitates near-real-time action prediction.
  • SR-VLN balances accuracy and efficiency, outperforming existing methods.