Related Experiment Videos
Toward Stronger 3D Human Action Representation Learning Via Efficient Spatiotemporal Decoupling
Abstract:
3D human action understanding has attracted extensive attention in both academia and industry, yet designing an effective representation learning paradigm remains a significant challenge. In this work, we present Stronger Spatiotemporal Decoupling Network (Stronger-SCD-Net), an effective self-supervised framework for 3D human action representation that systematically redesigns key components of the contrastive learning paradigm, including data augmentation, encoder architecture, and optimisation objective. To improve robustness and training efficiency, we propose a structurally constrained mask-out strategy based on deformation-driven data augmentation. This strategy not only accelerates the training process substantially but also improves the model's generalisation ability. For encoder design, we redesign the spatiotemporal decoupling mechanism into a more compact dual-branch architecture, in which spatial and temporal cues are explicitly modelled via factorised operations. This design improves representation discrimination and provides more interpretable spatial and temporal clues. In terms of the optimisation objective, we construct a global anchor representation to interactively align spatial and temporal cues from positive pairs while separating negative pairs. Furthermore, we incorporate an offline knowledge distillation strategy based on cosine similarity to further boost performance and stabilise representation learning under strong augmentation. Extensive experiments on NTU-RGB+D (60&120) and PKU-MMD (I&II) datasets demonstrate that the proposed Stronger-SCD-Net achieves superior performance across multiple downstream tasks, including action recognition, action retrieval, transfer learning, and semi-supervised learning. The code will be made available at https://github.com/cong-wu/SCD-Net.
Related Concept Videos
Three-Dimensional Force System:Problem Solving
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
Planar Rigid-Body Motion
Planar motion is typically divided into three distinct categories. The first is rectilinear translation, demonstrated by a subway train that moves along...
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it instrumental in...
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Maximizing the Directional Derivative
Divergence Theorem in 3D Space