Related Experiment Video
Updated: Sep 13, 2025

05:41
Author Spotlight: Integrating Ultrasound Imaging with Biochemical Markers for Thyroid Disease Diagnosis
Published on: February 9, 2024
740
DP-AMF: Depth-Prior-Guided Adaptive Multi-Modal and Global-Local Fusion for Single-View 3D Reconstruction
Luoxi Zhang1, Chun Xie2, Itaru Kitahara2
1Doctoral Program in Empowerment Informatics, University of Tsukuba, 1-1-1 Tennodai, Tsukuba 3058577, Japan.
Journal of Imaging
|July 25, 2025
Summary
This study introduces DP-AMF, a new framework for single-view 3D reconstruction. It effectively integrates depth priors with multi-modal features, improving geometric accuracy and detail in complex scenes.
Area of Science:
- Computer Vision
- 3D Computer Graphics
- Machine Learning
Background:
- Single-view 3D reconstruction is inherently challenging due to the lack of depth and scale information in 2D images.
- Ambiguities arise particularly in occluded or texture-poor areas, limiting reconstruction fidelity.
- Existing methods struggle to consistently recover accurate and detailed 3D geometry from a single image.
Purpose of the Study:
- To develop a robust framework for single-view 3D reconstruction that overcomes inherent ambiguities.
- To enhance the accuracy and completeness of reconstructed 3D models from single RGB images.
- To integrate diverse feature modalities effectively for improved geometric detail and handling of occlusions.
Main Methods:
- Proposed DP-AMF (Depth-Prior-Guided Adaptive Multi-Modal and Global-Local Fusion) framework.
- Integration of pre-computed, high-fidelity depth priors from a diffusion-based estimator (MARIGOLD).
- Fusion of hierarchical local features (ResNet) and semantic global features (DINO-ViT) with adaptive weighting.
- Utilized an implicit signed-distance field decoder for final mesh reconstruction.
Main Results:
- DP-AMF achieved significant improvements over strong baselines on 3D-FRONT and Pix3D datasets.
- Demonstrated a 7.64% reduction in Chamfer Distance, a 2.81% increase in F-Score, and a 5.88% boost in Normal Consistency.
- Qualitative results showcased sharper edges and more complete geometry, especially in challenging scenarios.
Conclusions:
- DP-AMF offers a robust and effective solution for complex single-view 3D reconstruction.
- The framework achieves state-of-the-art performance without substantial increases in model size or inference time.
- The adaptive multi-modal fusion approach effectively leverages depth priors and semantic features for superior reconstruction quality.

