Related Experiment Video
Updated: Sep 25, 2025

06:32
Author Spotlight: Automated Deep Brain Stimulation for Parkinson's Disease - Exploring the Possibilities and Challenges of Home Monitoring
Published on: July 14, 2023
1.5K
Dual Networks Based 3D Multi-Person Pose Estimation From Monocular Video.
Summary
This study introduces a novel approach for multi-person 3D human pose estimation by integrating top-down and bottom-up methods. The proposed technique enhances accuracy and robustness, especially in complex scenes with occlusions and interactions.
Area of Science:
- Computer Vision
- Machine Learning
- Robotics
Background:
- Monocular 3D human pose estimation is advancing, but current methods primarily focus on single individuals, using person-centric coordinates.
- These single-person methods are unsuitable for multi-person scenarios requiring absolute coordinates and struggle with challenges like occlusion and close interactions.
- Existing multi-person approaches (top-down and bottom-up) have limitations, including reliance on detection accuracy or errors with small-scale individuals.
Purpose of the Study:
- To develop a robust and accurate monocular multi-person 3D human pose estimation method.
- To overcome limitations of existing top-down and bottom-up approaches by integrating their strengths.
- To address challenges in multi-person pose estimation, such as occlusion, scale variation, and data scarcity.
Main Methods:
- Proposed an integrated network combining top-down and bottom-up approaches for multi-person 3D pose estimation.
- Developed a robust top-down network for estimating joints across all persons in image patches and a bottom-up network for scale variation resilience.
- Implemented test-time optimization using temporal constraints, re-projection loss, bone length regularizations, a two-person pose discriminator, and semi-supervised learning.
Main Results:
- The integrated approach demonstrated significant improvements in multi-person 3D human pose estimation accuracy and robustness.
- Individual components, including the enhanced top-down and bottom-up networks and the integration strategy, proved effective.
- The method successfully handled challenging scenarios involving inter-person occlusion and close interactions.
Conclusions:
- The proposed integrated method effectively addresses the limitations of existing single-person and multi-person 3D pose estimation techniques.
- The combination of top-down and bottom-up strategies, along with test-time optimization and semi-supervised learning, leads to superior performance.
- The developed approach offers a promising solution for accurate and reliable monocular multi-person 3D pose estimation in complex real-world scenarios.

