Related Experiment Video
Updated: Jun 7, 2026

Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
Published on: April 21, 2023
Learning multi-granularity skeleton representations via hierarchical graph contrastive learning
Yaxiong Liu1, Xianfeng Zhai2,3
1Qinghai Normal University, Qinghai, 810000, China.
Abstract:
Skeleton-based action recognition offers inherent robustness to illumination variations and background interference, yet faces challenges such as heavy label reliance, topological destruction during augmentation, and underutilized joint confidence. This paper proposes Hierarchical Graph Contrastive Learning (HGCL), in which the three challenges are addressed in a coupled manner rather than through independently inserted modules. The augmentation strategy combines joint masking, spatial rotation, and Gaussian jittering under a unified topology constraint, with perturbation intensity adapted to the graph distance from core joints and a provable upper bound on bone length deviation. The hierarchical contrastive objective decomposes the InfoNCE loss into a global sequence branch and a part-level branch in which negative pairs are restricted within the same body region, so that local motion signatures and holistic action semantics are jointly optimized rather than entangled at a single granularity. The detection confidence is propagated through both the adjacency matrix used for graph convolution and the pooling operator used for representation aggregation, providing a single mechanism that suppresses unreliable joints throughout the pipeline. Experiments on UCF101 and HMDB51 show that HGCL achieves 91.2% and 66.8% Top-1 accuracy respectively, outperforming InfoGCN by 3.9 and 5.3 percentage points. In 1% few-shot scenarios, HGCL improves upon training-from-scratch by 21.6 percentage points on UCF101. Controlled degradation experiments further demonstrate that the performance advantage of HGCL over baselines widens under increasing coordinate noise and joint occlusion. Feature space analysis reveals a 54.5% reduction in intra-class distance and a 90.3% increase in inter-class distance, validating its effectiveness for self-supervised representation learning.
Related Concept Videos
Classification of Bones
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The long...
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting the...
Observational Learning
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Associative Learning
Classical conditioning, also known...
Structural Classification of Joints
A fibrous joint is where the adjacent bones are united by fibrous connective...