Related Experiment Video
Updated: Jun 4, 2025

09:41
Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
Published on: April 21, 2023
1.5K
Deep Fusion of Skeleton Spatial-Temporal and Dynamic Information for Action Recognition
Song Gao1, Dingzhuo Zhang2, Zhaoming Tang1
1Aviation Maintenance NCO Academy, Air Force Engineering University, Xinyang 464007, China.
Sensors (Basel, Switzerland)
|December 17, 2024
Summary
This study introduces a novel action recognition approach using skeleton spatial-temporal and dynamic features with a two-stream convolutional neural network (TS-CNN). The method significantly improves human action recognition accuracy in complex environments.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Traditional deep learning methods for action recognition suffer from low recognition rates.
- Existing algorithms often fail to capture complex human motion effectively.
Purpose of the Study:
- To develop an improved action recognition algorithm using skeleton spatial-temporal and dynamic features.
- To enhance the accuracy and robustness of human action recognition in complex environments.
Main Methods:
- Skeleton data was transformed to extract relative joint positions and encoded as color texture maps for spatial-temporal features.
- Human body constraints were used to improve inter-class distinctions.
- Joint speed information was encoded as color texture maps for dynamic features.
- Features were enhanced using motion saliency and morphology operators.
- A two-stream convolutional neural network (TS-CNN) was employed for deep fusion and action recognition.
Main Results:
- The developed approach achieved high recognition rates: 86.25% on NTU RGB-D, 87.37% on Northwestern-UCLA, and 93.75% on UTD-MHAD.
- The method demonstrated superior performance compared to state-of-the-art algorithms.
Conclusions:
- The proposed skeleton-based action recognition method effectively improves accuracy.
- The combination of spatial-temporal, dynamic features, and TS-CNN offers a robust solution for complex human action recognition.

