Related Experiment Video
Updated: May 23, 2026

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
AstroHSP: A hybrid supervision framework for robust monocular astronaut pose estimation
Haohang Jian1, Yuhao Xiao2, Xiongwu Xiao1
1State Key Laboratory of Information Engineering in Surveying, Mapping and Remote sensing, Wuhan University, Wuhan, Hubei, China.
Abstract:
Accurate 3D human pose estimation is critical for astronaut safety and ergonomic analysis, yet it remains a fundamental challenge under severe deformation, heavy self-occlusion, and extreme domain shift, particularly in the absence of 3D ground-truth annotations. To address this problem, we propose AstroHSP, a robust two-stage framework that integrates domain-adaptive 2D pose estimation with uncertainty-aware 3D pose lifting. In the first stage, a mixed-domain training strategy combined with a novel Domain-Adaptive Batch Normalization mechanism is adopted to achieve stable and consistent 2D keypoint predictions under extreme appearance distortions and severe occlusions. In the second stage, we design an advanced 2D-to-3D lifting module that employs a dual-stream Transformer to model spatio-temporal dependencies and incorporates a conditional diffusion model to iteratively refine coarse 3D poses, thereby effectively modeling and mitigating inherent depth uncertainty. To cope with the practical unavailability of 3D ground truth in highly constrained target domains, we further introduce a two-stage hybrid training strategy, in which the model is first pre-trained with full supervision on public datasets and subsequently fine-tuned through joint optimization that retains full supervision on public data while introducing weak geometric constraints on the target domain. We introduce AstPose, a dataset collected from real space missions, and conduct comprehensive evaluations across multiple datasets, including clothing-heavy and highly constrained scenarios. AstroHSP demonstrates superior performance across multiple datasets, affirming the efficacy of our specialized design for robust 3D pose generalization in highly constrained domains.

