Related Experiment Video
Updated: Aug 6, 2026

Deep-Learning Based Multi-Joint Synchronous Tracking for Objective Quantification of Hindlimb Locomotor Kinematics in Rats
Published on: April 3, 2026
Comparative evaluation of deep learning models for three-class frailty assessment using gait metrics
Charmayne Mary Lee Hughes1, Yan Zhang1
1Age-Appropriate Human-Machine Systems, Institute of Psychology and Ergonomics, Technische Universität Berlin, Berlin, Germany.
Background:
Frailty assessment in older adults typically relies on subjective clinical tools that are time-consuming and require trained personnel, limiting their use in routine or large-scale screening. Wearable sensor-based gait analysis offers a promising objective alternative; however, comprehensive evaluations of deep learning models for multi-class frailty classification using structured gait metrics remain limited.
Methods:
This study evaluates multiple deep learning architectures for three-class frailty classification using structured gait features derived from a single wearable inertial measurement unit (IMU). Four architectures (Transformer, ShapeFormer, InceptionTime, and LSTM-CNN) were evaluated using the publicly available GSTRIDE dataset, comprising 163 older adults (non-frail: n = 65, pre-frail: n = 77, frail: n = 21). Eleven clinically interpretable stride-level gait variables, extracted from a single IMU, were used as model input. A 10-fold participant-level cross-validation scheme, matching the default configuration of the original implementation (cv_seed = 1), was applied to ensure robust performance estimation and prevent data leakage.
Results:
The four architectures yielded comparable mean accuracy under a unified protocol (ShapeFormer 73.1% ± 10.6%, Transformer 72.6% ± 6.0%, InceptionTime 70.5% ± 9.5%, LSTM-CNN 65.9% ± 9.8%). A Friedman omnibus test on the per-fold scores did not reject the null of equal performance (borderline on accuracy and macro-F1 (p = 0.050 and 0.056 respectively), with no significant difference for macro-AUC [p = 0.840]), indicating that the apparent ranking is within sampling variability. Performance was higher for non-frail and pre-frail individuals, while classification of frail individuals remained challenging across all models.
Conclusion:
Multi-class frailty classification using structured gait metrics from a single wearable sensor is feasible, though performance varies across models and frailty classes. Model stability and sensitivity to specific classes represent important trade-offs.
Significance:
These findings support the development of scalable, objective approaches to frailty assessment and underscore the importance of selecting models based on clinical priorities.

