Related Experiment Video
Updated: Dec 31, 2025

Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
Cycle-SfM: Joint self-supervised learning of depth and camera motion from monocular image sequences
Qiyu Sun1, Yang Tang1, Chaoqiang Zhao1
1The Key Laboratory of Advanced Control and Optimization for Chemical Processes, Ministry of Education, East China University of Science and Technology, Shanghai 200237, China.
This study introduces a self-supervised framework for estimating 3D scene geometry, specifically monocular depth and camera ego-motion, from unlabelled videos. The method improves accuracy through a novel cost function and forward-backward consistency.
Area of Science:
- Computer Vision
- Robotics
- Artificial Intelligence
Background:
- 3D scene geometry understanding is crucial for computer vision tasks like depth prediction and visual odometry.
- Deep learning enables end-to-end solutions by framing 3D understanding as nonlinear optimization problems.
- Existing methods often require labeled data or complex pipelines.
Purpose of the Study:
- To develop a self-supervised framework for joint monocular depth and camera ego-motion estimation.
- To leverage unlabeled, unstructured monocular video sequences for 3D scene understanding.
- To improve the accuracy and efficiency of 3D geometry estimation.
Main Methods:
- A self-supervised framework utilizing forward-backward consistency constraint on view reconstruction.
- Exploiting bidirectional projection information across adjacent video frames.
- Introducing a lightweight and generalizable cost function improvement for enhanced accuracy.
Main Results:
- The proposed framework achieves comparable depth estimation results to existing methods.
- Demonstrates superior performance in pose estimation compared to current approaches on the KITTI dataset.
- The cost function improvement module is shown to be effective and seamlessly integrable.
Conclusions:
- The developed self-supervised method effectively estimates monocular depth and camera ego-motion from unlabelled video.
- The novel cost function enhancement offers a practical and efficient way to boost accuracy in self-supervised 3D vision tasks.
- The framework presents a significant advancement in efficient and accurate 3D scene geometry understanding.
Related Concept Videos
Depth Perception and Spatial Vision
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
Absolute Motion Analysis- General Plane Motion
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the...
Observational Learning
Uniform Depth Channel Flow: Problem Solving
