Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Visual System01:26

Visual System

645
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
645
Vision01:24

Vision

54.8K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
54.8K
Parallel Processing01:20

Parallel Processing

203
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
203
Visual Agnosia01:12

Visual Agnosia

268
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
268
Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

816
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
816
Higher Mental Functions of the Brain: Language01:10

Higher Mental Functions of the Brain: Language

967
Language is a system of communication that allows the expression of thoughts, ideas, and feelings. The brain processes language in both hemispheres.
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
967

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Dual-branch underwater image enhancement network based on fusion of polarization andĀ colorĀ information.

Applied opticsĀ·2026
Same author

Polarization-guided diffusion model for physically inspired underwater image descattering.

Applied opticsĀ·2026
Same author

PRAFNet: a polarization-RGB adaptive fusion network for object detection in complex environments.

Applied opticsĀ·2025
Same author

Polarization-driven camouflaged object detection: a multimodal fusion network with iterative polarimetric feature enhancement.

Applied opticsĀ·2025
Same author

Research on a hybrid neural network task assignment algorithm for solving multi-constraint heterogeneous autonomous underwater robot swarms.

Frontiers in neuroroboticsĀ·2023
Same author

A Transcriptome-Wide Association Study Identifies Novel Candidate Susceptibility Genes for Pancreatic Cancer.

Journal of the National Cancer InstituteĀ·2020

Related Experiment Video

Updated: Aug 19, 2025

Development of an Audio-based Virtual Gaming Environment to Assist with Navigation Skills in the Blind
09:01

Development of an Audio-based Virtual Gaming Environment to Assist with Navigation Skills in the Blind

Published on: March 27, 2013

14.5K

Vital information matching in vision-and-language navigation.

Zixi Jia1, Kai Yu1, Jingyu Ru1

  • 1Faculty of Robot Science and Engineering, Northeastern University, Shenyang, China.

Frontiers in Neurorobotics
|December 5, 2022
PubMed
Summary

Researchers developed a new AI model, the Vital Information Matching Feedback Self-tuning Network (VIM-Net), to improve visual language navigation by better fusing multi-modal inputs. This novel approach enhances how AI understands and navigates environments based on visual and textual cues.

Keywords:
collaborative learningmultimodal matchingself-tuning modulevision-and-language navigationvital information matching networks

More Related Videos

Practical Methodology of Cognitive Tasks Within a Navigational Assessment
05:19

Practical Methodology of Cognitive Tasks Within a Navigational Assessment

Published on: June 1, 2015

13.7K
Author Spotlight: Investigating the Effects of Mind-Body-Movement Practices on Brain Function
06:17

Author Spotlight: Investigating the Effects of Mind-Body-Movement Practices on Brain Function

Published on: January 26, 2024

2.1K

Related Experiment Videos

Last Updated: Aug 19, 2025

Development of an Audio-based Virtual Gaming Environment to Assist with Navigation Skills in the Blind
09:01

Development of an Audio-based Virtual Gaming Environment to Assist with Navigation Skills in the Blind

Published on: March 27, 2013

14.5K
Practical Methodology of Cognitive Tasks Within a Navigational Assessment
05:19

Practical Methodology of Cognitive Tasks Within a Navigational Assessment

Published on: June 1, 2015

13.7K
Author Spotlight: Investigating the Effects of Mind-Body-Movement Practices on Brain Function
06:17

Author Spotlight: Investigating the Effects of Mind-Body-Movement Practices on Brain Function

Published on: January 26, 2024

2.1K

Area of Science:

  • Artificial Intelligence
  • Multi-modal Machine Learning
  • Robotics

Background:

  • Visual language navigation is a key task in multi-modal machine learning, crucial for integrating information from different input types.
  • Current models struggle to effectively fuse multi-modal inputs, limiting their ability to capture intrinsic relationships between data sources.
  • Existing methods often rely on basic data augmentation, failing to fully exploit the potential of multi-modal interactions.

Purpose of the Study:

  • To propose a novel multi-modal matching feedback self-tuning model to address the limitations of existing visual language navigation systems.
  • To introduce the Vital Information Matching Feedback Self-tuning Network (VIM-Net) for enhanced fusion of visual and textual information.
  • To improve the performance of AI agents in understanding and executing navigation commands within complex environments.

Main Methods:

  • Developed the Vital Information Matching Feedback Self-tuning Network (VIM-Net), a novel neural network architecture.
  • Implemented two core matching feedback modules: a visual matching feedback module (V-mat) and a trajectory matching feedback module (T-mat).
  • V-mat aligns visual recognition targets with command-extracted entity information; T-mat matches serialized trajectory features with command-specified movement directions.

Main Results:

  • Conducted ablation and comparative experiments using the Matterport3D simulator and Room-to-Room (R2R) benchmark datasets.
  • Demonstrated the effectiveness of VIM-Net through detailed analysis of navigation outcomes.
  • The proposed model achieved significant improvements in visual language navigation tasks.

Conclusions:

  • The novel VIM-Net model effectively addresses the challenge of fusing multi-modal inputs for visual language navigation.
  • The proposed matching feedback modules (V-mat and T-mat) are crucial for enhancing the model's performance.
  • Experimental results validate the efficacy of VIM-Net on benchmark datasets, proving its practical applicability.