Related Experiment Video
Updated: Sep 9, 2026

Deep-Learning Based Multi-Joint Synchronous Tracking for Objective Quantification of Hindlimb Locomotor Kinematics in Rats
Published on: April 3, 2026
SE-ResAutoNet: an attention and residual learning framework for reliable gait-based suspect identification
Vaishnavi Munusamy1, Sudha Senthilkumar1
1School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, Tamil Nadu, India.
Abstract:
Gait recognition relies on the distinctive walking pattern of the individuals that helps to identify the person accurately even in the long distance or in varied environments without any user interaction. Unlike existing techniques, the method works independently, it does not depend on a person engaging or interacting directly with the system. The SE-ResAutoNet (Squeeze-and-Excitation Residual Autoencoder Network) proposed in this study, an autoencoder framework which adopts SE blocks and Residual Blocks in their architecture for better feature representation and classification accuracy. The attention based - residual learning helps to capture the complex pattern effectively. The SE-ResAutoNet model has achieved an average accuracy of 95.75% across 5-fold cross-validation under normal, slow, and fast walking conditions, along with high precision, recall, and Rank-1 identification performance, demonstrating the consistency and robustness of the model. Furthermore, the proposed framework achieved benchmark accuracies of 97.1% on CASIA-B and 96.32% on OU-MVLP datasets, indicating strong cross-dataset generalization capability. The model also demonstrated computational efficiency with a minimal inference time of 11 ms per frame (77 FPS) and a compact computational cost of 1.5 GFLOPs, highlighting its suitability for real-world video-based surveillance applications. In addition, inclusion of Integrated Gradients highlights the most influential gait regions in the frame. The finding suggests that SE-ResAutoNet achieves an effective balance between accuracy, computational efficiency, and explainability, highlighting its strong potential for deployment in practical biometric and surveillance applications, subject to further validation on larger and more diverse datasets.

