Related Experiment Video
Updated: Oct 25, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
7.9K
Unsupervised Estimation of Monocular Depth and VO in Dynamic Environments via Hybrid Masks
IEEE Transactions on Neural Networks and Learning Systems
|August 4, 2021
Summary
This study introduces hybrid masks to improve deep learning for 3-D sensing, enhancing depth and visual odometry (VO) estimation in dynamic environments. The method achieves state-of-the-art results on benchmarks.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Robotics
Background:
- Deep learning excels in 3-D sensing but struggles with dynamic environments.
- Existing monocular visual odometry (VO) methods are susceptible to failures caused by moving objects.
Purpose of the Study:
- To mitigate the negative impact of dynamic environments on joint depth and VO estimation.
- To improve the robustness and accuracy of 3-D perception systems.
Main Methods:
- Proposed hybrid masks (cover and filter masks) to address dynamic elements in VO and depth estimation.
- Introduced a depth-pose consistency loss to resolve scale inconsistencies in monocular sequences.
- Developed a joint estimation framework for depth and VO.
Main Results:
- Achieved state-of-the-art depth prediction and globally consistent VO estimation on the KITTI benchmark.
- Demonstrated the method's transferability on the Make3D dataset.
- Hybrid masks effectively reduced the adverse effects of dynamic environments.
Conclusions:
- The proposed hybrid mask approach significantly enhances deep learning-based 3-D sensing in dynamic scenarios.
- The method provides accurate and robust depth and VO estimation, crucial for autonomous systems.
- The depth-pose consistency loss further improves the reliability of monocular sequence analysis.
More Related Videos
Related Concept Videos
Depth Perception and Spatial Vision
1.2K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.2K
Masking and Demasking Agents
2.9K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.9K
Relative Motion Analysis using Rotating Axes-Problem Solving
495
Consider a crane whose telescopic boom rotates with an angular velocity of 0.04 rad/s and angular acceleration of 0.02 rad/s2. Along with the rotation, the boom also extends linearly with a uniform speed of 5 m/s. The extension of the boom is measured at point D, which is measured with respect to the fixed point C on the other end of the boom. For the given instant, the distance between points C and D is 60 meters.
Here, in order to determine the magnitude of velocity and acceleration for point...
Here, in order to determine the magnitude of velocity and acceleration for point...
495
Relative Motion Analysis using Rotating Axes
611
Consider a component AB undergoing a linear motion. Along with a linear motion, point B also rotates around point A. To comprehend this complex movement, position vectors for both points A and B are established using a stationary reference frame.
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
611

