Related Experiment Video
Updated: Nov 18, 2025

05:55
Modeling the Functional Network for Spatial Navigation in the Human Brain
Published on: October 13, 2023
1.3K
Joint Multimodal Embedding and Backtracking Search in Vision-and-Language Navigation
1Department of Computer Science, Kyonggi University, Suwon-si 16227, Korea.
Sensors (Basel, Switzerland)
|February 5, 2021
Summary
This study introduces JMEBS, a novel deep neural network for vision-and-language navigation (VLN). It enhances navigation success rates and path optimization using joint multimodal embedding and backtracking search.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Multimodal intelligent tasks integrating vision and language are gaining prominence.
- Vision-and-language navigation (VLN) requires aligning and grounding image and text data for real-time task status perception.
Purpose of the Study:
- To propose a novel deep neural network model, JMEBS, for enhanced performance in vision-and-language navigation tasks.
- To improve task success rates and optimize navigation paths through advanced embedding and search algorithms.
Main Methods:
- Developed a transformer-based joint multimodal embedding module for JMEBS, utilizing both multimodal and temporal contexts.
- Implemented backtracking-enabled greedy local search (BGLS) with a novel global scoring method for action selection and trajectory evaluation.
- Evaluated the model using the Matterport3D Simulator and room-to-room (R2R) benchmark datasets.
Main Results:
- The JMEBS model demonstrated improved task success rates and optimized navigation paths compared to existing models.
- The novel global scoring method effectively improved performance by comparing partial trajectories with natural language instructions.
- Experimental results validated the model's effectiveness across various operations on benchmark datasets.
Conclusions:
- The proposed JMEBS model offers a significant advancement in vision-and-language navigation.
- The integration of joint multimodal embedding and backtracking search effectively addresses key challenges in VLN.
- JMEBS provides a robust framework for future research in embodied AI and multimodal understanding.
Related Concept Videos
Depth Perception and Spatial Vision
1.4K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.4K
Collisions in Multiple Dimensions: Problem Solving
4.8K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.8K
Visual Agnosia
639
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
639
Vision
58.5K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
58.5K
Observational Learning
617
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
617
Associative Learning
862
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
862

