Related Experiment Video
Updated: Jun 24, 2025

06:25
Time Multiplexing Super Resolving Technique for Imaging from a Moving Platform
Published on: February 12, 2014
8.5K
SSR-2D: Semantic 3D Scene Reconstruction From 2D Images
Summary
This study introduces a novel deep learning method for 3D indoor scene reconstruction and semantic mapping using only 2D labels, eliminating the need for expensive 3D annotations. The approach achieves state-of-the-art results on benchmark datasets, enabling efficient semantic scene completion.
Area of Science:
- Computer Vision
- 3D Scene Understanding
- Machine Learning
Background:
- Deep learning for 3D indoor scene modeling typically requires extensive 3D annotations, which are costly and time-consuming to obtain.
- Existing methods struggle with comprehensive semantic modeling due to annotation limitations.
Purpose of the Study:
- To develop a novel approach for semantic scene reconstruction in 3D indoor spaces that eliminates the need for 3D annotations.
- To enable joint geometry completion, colorization, and semantic mapping using only 2D labels and RGB-D images.
Main Methods:
- A trainable model fusing cross-domain features from incomplete 3D reconstructions and RGB-D images into volumetric embeddings.
- Leveraging differentiable rendering to bridge 2D observations (RGB images and 2D semantics) with unknown 3D space for supervision.
- Developing a learning pipeline for self-supervision using imperfect predicted 2D labels from synthesized virtual views.
Main Results:
- Achieved state-of-the-art performance on semantic scene completion tasks on MatterPort3D and ScanNet datasets.
- Outperformed baseline methods that utilized costly 3D annotations for geometry and semantic prediction.
- Demonstrated simultaneous completion and semantic segmentation of real-world 3D scans using a 2D-driven approach.
Conclusions:
- The proposed 2D-driven method offers an efficient and effective solution for 3D semantic scene reconstruction without 3D ground truth.
- This approach significantly reduces the annotation burden, making comprehensive 3D scene modeling more accessible.
- The method represents a significant advancement in 2D-driven 3D scene understanding and reconstruction.

