Related Experiment Video
Updated: Dec 15, 2025

Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
Joint Unsupervised Learning of Depth, Pose, Ground Normal Vector and Ground Segmentation by a Monocular Camera Sensor
Lu Xiong1, Yongkun Wen1, Yuyao Huang1
1Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China.
Abstract:
We propose a completely unsupervised approach to simultaneously estimate scene depth, ego-pose, ground segmentation and ground normal vector from only monocular RGB video sequences. In our approach, estimation for different scene structures can mutually benefit each other by the joint optimization. Specifically, we use the mutual information loss to pre-train the ground segmentation network and before adding the corresponding self-learning label obtained by a geometric method. By using the static nature of the ground and its normal vector, the scene depth and ego-motion can be efficiently learned by the self-supervised learning procedure. Extensive experimental results on both Cityscapes and KITTI benchmark demonstrate the significant improvement on the estimation accuracy for both scene depth and ego-pose by our approach. We also achieve an average error of about 3° for estimated ground normal vectors. By deploying our proposed geometric constraints, the IOUaccuracy of unsupervised ground segmentation is increased by 35% on the Cityscapes dataset.
Related Concept Videos
Depth Perception and Spatial Vision
Uniform Depth Channel Flow: Problem Solving
Curvilinear Motion: Normal and Tangential Components
The positive direction of the t-axis aligns with the increasing position of the car along the curved path, denoted by the unit vector ut. Simultaneously, the n-axis, perpendicular to the t-axis, dissects the curved path into differential arc segments, each forming the arc of a circle with a radius of...
