GaitTriViT and GaitVViT: Transformer-based methods emphasizing spatial or temporal aspects in gait recognition.
1School of Computing Science, University of Glasgow, Glasgow, United Kingdom.
Peerj. Computer Science
|September 24, 2025
Summary
This study explores Transformer-based methods for gait recognition, a biometric technology. While Vision Transformers show promise, challenges remain in achieving state-of-the-art performance for identifying individuals by their walking patterns.
Area of Science:
- Computer Vision
- Biometrics
- Artificial Intelligence
Background:
- Gait recognition is a promising biometric technology, but challenges persist with low-resolution and long-distance subject identification.
- Traditional methods often rely on Convolutional Neural Networks (CNNs), which may not fully capture complex spatio-temporal gait features.
- Vision Transformers (ViTs) have emerged as powerful alternatives in computer vision, offering potential for improved performance.
Purpose of the Study:
- To introduce and evaluate two novel Transformer-based methods for gait recognition: GaitTriViT and GaitVViT.
- To assess the effectiveness of ViTs in extracting fine-grained spatial and temporal features for gait analysis.
- To compare the performance of these new methods against current state-of-the-art (SOTA) approaches.
Main Methods:
- Developed GaitTriViT, a method utilizing Vision Transformers (ViTs) for enhanced spatial feature extraction.
- Developed GaitVViT, a method employing Video Vision Transformers (Video ViT) to improve temporal feature understanding.
- Evaluated the performance of both methods on gait recognition tasks, comparing them with existing SOTA models.
Main Results:
- The proposed Transformer-based methods, GaitTriViT and GaitVViT, demonstrate encouraging performance in gait recognition.
- Results indicate existing gaps and challenges in applying ViTs to gait recognition, particularly with complex scenarios.
- The study highlights both the potential and the difficulties encountered when using ViTs for this biometric task.
Conclusions:
- Transformer-based methods, specifically ViTs, hold significant promise for advancing gait recognition technology.
- Further research and development are needed to overcome the identified challenges and fully leverage the capabilities of ViTs in this domain.
- The future of Vision Transformers in gait recognition appears promising, despite current limitations.


