Related Experiment Video
Updated: Dec 15, 2025

Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
Joint Unsupervised Learning of Depth, Pose, Ground Normal Vector and Ground Segmentation by a Monocular Camera Sensor
Lu Xiong1, Yongkun Wen1, Yuyao Huang1
1Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China.
This study introduces an unsupervised method for estimating scene depth, ego-pose, ground segmentation, and ground normal vectors from monocular video. The approach enhances accuracy for depth and ego-pose estimation and improves ground segmentation by 35%.
Area of Science:
- Computer Vision
- Robotics
- Machine Learning
Background:
- Autonomous systems require accurate perception of scene geometry and self-motion.
- Existing methods often rely on supervised learning or multiple sensors, limiting their applicability.
- Unsupervised learning offers a promising avenue for robust perception from monocular video.
Purpose of the Study:
- To develop a completely unsupervised approach for simultaneous estimation of scene depth, ego-pose, ground segmentation, and ground normal vectors.
- To leverage the mutual benefits of joint optimization for different scene structure estimations.
- To improve the accuracy and robustness of perception tasks in monocular video sequences.
Main Methods:
- Utilizing a joint optimization framework for simultaneous estimation of multiple scene properties.
- Employing mutual information loss for pre-training the ground segmentation network.
- Incorporating self-learning labels derived from geometric methods for ground segmentation.
- Leveraging the static nature of ground and its normal vector for self-supervised depth and ego-motion learning.
Main Results:
- Significant improvements in estimation accuracy for scene depth and ego-pose on Cityscapes and KITTI benchmarks.
- Achieved an average error of approximately 3° for estimated ground normal vectors.
- Increased the Intersection over Union (IOU) accuracy of unsupervised ground segmentation by 35% on the Cityscapes dataset.
Conclusions:
- The proposed unsupervised approach effectively estimates multiple scene properties from monocular video.
- Joint optimization and geometric constraints significantly enhance perception accuracy and robustness.
- This method offers a viable solution for perception in resource-constrained or unsupervised scenarios.
Related Concept Videos
Depth Perception and Spatial Vision
Uniform Depth Channel Flow: Problem Solving
Curvilinear Motion: Normal and Tangential Components
The positive direction of the t-axis aligns with the increasing position of the car along the curved path, denoted by the unit vector ut. Simultaneously, the n-axis, perpendicular to the t-axis, dissects the curved path into differential arc segments, each forming the arc of a circle with a radius of...
