Related Experiment Video
Updated: Nov 24, 2025

09:41
Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
Published on: April 21, 2023
2.0K
Masked GAN for Unsupervised Depth and Pose Prediction With Scale Consistency
IEEE Transactions on Neural Networks and Learning Systems
|December 28, 2020
Summary
This study introduces a masked generative adversarial network (GAN) for improved unsupervised monocular depth and ego-motion estimation. The novel approach effectively handles occlusions and visual field changes, enhancing camera trajectory prediction.
Area of Science:
- Computer Vision
- Machine Learning
- Robotics
Background:
- Unsupervised monocular depth and visual odometry (VO) estimation commonly use adversarial learning with reconstruction losses.
- Performance is often hindered by occlusions and changing visual fields between frames.
Purpose of the Study:
- To propose a masked generative adversarial network (GAN) for robust unsupervised monocular depth and ego-motion estimation.
- To mitigate the impact of occlusions and visual field variations on estimation accuracy.
Main Methods:
- Introduced MaskNet and a Boolean mask scheme to filter out occluded or changed regions.
- Implemented a scale-consistency loss for accurate long-term camera trajectory estimation.
- Utilized adversarial and geometric image reconstruction losses as primary training signals.
Main Results:
- Demonstrated that each proposed component improves performance.
- Achieved competitive results for both depth and trajectory predictions on KITTI and Make3D datasets.
- The masking strategy effectively reduced negative impacts from occlusions and visual field changes.
Conclusions:
- The masked GAN framework offers a significant advancement in unsupervised monocular depth and ego-motion estimation.
- The proposed methods provide a more accurate and robust solution for monocular sequence analysis.
- This work paves the way for more reliable visual odometry and depth perception in challenging environments.

