Related Experiment Video
Updated: Jul 15, 2025

09:10
Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
1.8K
Dynamic surface reconstruction in robot-assisted minimally invasive surgery based on neural radiance fields
Xinan Sun1,2, Feng Wang1, Zhikang Ma1
1School of Mechanical Engineering, Tianjin University, 135 Yaguan Road, Jinnan District, Tianjin, 300350, China.
International Journal of Computer Assisted Radiology and Surgery
|September 28, 2023
Summary
This study introduces a new depth estimation network and neural radiance fields framework to improve surgical scene perception and dynamic scene reconstruction for better surgical automation and AR navigation.
Area of Science:
- Computer Vision
- Medical Imaging
- Robotics
Background:
- Surgical scene perception is crucial for task automation and AR navigation.
- Reconstructing highly dynamic surgical scenes presents significant challenges.
- Existing methods struggle with accuracy and deformation in complex surgical environments.
Purpose of the Study:
- To enhance surgical scene perception by developing a novel depth estimation network.
- To create a robust reconstruction framework for dynamic surgical scenes using neural radiance fields.
- To provide more accurate scene information for surgical task automation and augmented reality (AR) navigation.
Main Methods:
- Implemented a depth estimation network with spatial pyramid pooling and Swin-Transformer modules for enhanced robustness.
- Incorporated optimal transport matching constraints to improve depth accuracy.
- Utilized neural radiance fields for implicit scene representation in the time dimension, optimized with depth and color information to prevent deformation distortion.
Main Results:
- The depth estimation network achieved near state-of-the-art (SOTA) performance on natural images and surpassed SOTA on medical images (1.12% in 3 px Error, 0.45 px EPE).
- The dynamic reconstruction framework successfully reconstructed cardiac surfaces from endoscopic videos, achieving SOTA metrics (27.983 dB PSNR, 0.812 SSIM, 0.189 LPIPS).
- Experiments validated performance on KITTI, SCARED, and real surgical datasets.
Conclusions:
- The proposed depth estimation network and reconstruction framework significantly advance surgical scene perception.
- The framework yields more accurate depth maps with clearer edges compared to SOTA methods on medical datasets.
- The framework was successfully verified on dynamic cardiac surgical images, with future work focusing on training speed and field-of-view limitations.

