Related Experiment Video
Updated: Jan 12, 2026

05:12
Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
2.4K
Consistent 3D Human Reconstruction From Monocular Video: Learning Correctable Appearance and Temporal Motion Priors.
IEEE Transactions on Visualization and Computer Graphics
|October 31, 2025
Summary
This study introduces novel methods for rendering dynamic humans from monocular video, improving motion realism and frame continuity. The approach enhances digital human representation for more accurate viewpoint rendering.
Area of Science:
- Computer Vision
- Computer Graphics
- Machine Learning
Background:
- Recent progress in rendering dynamic humans using Neural Radiance Fields (NeRF) and 3D Gaussian Splatting has advanced digital human creation.
- Monocular video rendering faces challenges with viewpoint imbalance, affecting subtle motion rendering and temporal continuity from novel viewpoints.
Purpose of the Study:
- To address limitations in monocular dynamic human rendering, specifically improving motion subtlety and inter-frame continuity.
- To enhance the accuracy and realism of digital humans rendered from various perspectives.
Main Methods:
- A pixel-level motion correction module was developed to rectify representation errors across different viewpoints.
- A temporal information-based model was introduced to enhance motion continuity by utilizing adjacent video frames.
Main Results:
- The proposed method demonstrates superior performance in dynamic human rendering compared to existing state-of-the-art techniques.
- Quantitative and qualitative evaluations on benchmark datasets (NeuMan, ZJU-Mocap, People-Snapshot) confirm the effectiveness of the approach.
Conclusions:
- The developed techniques significantly improve the rendering of dynamic humans from monocular video, particularly in handling complex motion and maintaining temporal consistency.
- This work advances the capabilities of creating realistic and continuous digital humans from limited visual input.

