Related Experiment Video
Updated: May 28, 2025

Estimation of Contact Regions Between Hands and Objects During Human Multi-Digit Grasping
Published on: April 21, 2023
Spatio-Temporal Transformer with Kolmogorov-Arnold Network for Skeleton-Based Hand Gesture Recognition
Pengcheng Han1, Xin He1, Takafumi Matsumaru1
1Graduate School of Information, Production and System, Waseda University, Kitakyushu 808-0135, Japan.
This study introduces the ST-KT framework for skeleton-based hand gesture recognition, utilizing spatio-temporal graph convolutions and a transformer with Kolmogorov-Arnold Networks (KAN) to capture complex joint dynamics for improved accuracy.
Area of Science:
- Computer Vision
- Machine Learning
- Human-Computer Interaction
Background:
- Manual feature engineering for gesture recognition is subjective and lacks robustness.
- Existing deep learning models often neglect crucial spatial-temporal and structural hand joint information.
- Capturing long-range dependencies between non-adjacent hand joints is vital for accurate recognition.
Purpose of the Study:
- To propose an advanced skeleton-based hand gesture recognition framework, ST-KT.
- To effectively model both spatial and temporal dependencies within human hand joint data.
- To leverage the strengths of graph convolutional networks and transformer architectures enhanced with Kolmogorov-Arnold Networks (KAN).
Main Methods:
- The ST-KT framework integrates spatio-temporal graph convolution network (ST-GCN) modules and a KAN-based transformer.
- ST-GCN modules (comprising spatial graph convolution network and temporal convolution network) extract initial skeleton sequence features.
- A spatio-temporal position embedding method enriches node representations with identity and temporal context, while KAN-Transformers capture intricate joint relationships.
Main Results:
- The proposed ST-KT method achieved high accuracy on challenging datasets: 97.5% on SHREC'17 and 94.3% on DHG-14/28.
- The framework effectively captures dynamic skeleton changes and complex inter-joint relationships.
- The integration of KAN within the transformer enhanced nonlinear modeling capabilities for richer feature extraction.
Conclusions:
- The ST-KT framework demonstrates superior performance in skeleton-based dynamic hand gesture recognition.
- The method successfully addresses limitations of previous approaches by incorporating spatio-temporal dynamics and long-range joint dependencies.
- This research offers a robust and accurate solution for advanced human-computer interaction applications.
More Related Videos
Related Concept Videos
Sequence Networks of Rotating Machines
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
Kinematic Equations for Rotation
For instance, imagine a point A on a rigid body engaged in circular motion. The translational velocity of this particular point can be calculated by taking the time derivatives of the displacement equation, which essentially measures the...

