Related Experiment Video
Updated: Aug 28, 2026

Dynamic Digital Biomarkers of Motor and Cognitive Function in Parkinson's Disease
Published on: July 24, 2019
Multi-Task Wearable Parkinson's Disease Detection with a Pretrained Spatio-Temporal Graph Encoder and Task-Level
H M K K M B Herath1, Nuwan Madusanka2, Chaminda Hewage3
1Industry 4.0 Convergence Bionics Engineering, Pukyong National University, Busan 48513, Republic of Korea.
Abstract:
Wearable inertial measurement units (IMUs) offer an objective, low-cost basis for Parkinson's disease (PD) assessment, but multi-task clinical protocols yield heterogeneous recordings across body locations and small cohorts, and it is unclear whether such data can support reliable PD detection without training deep models from scratch. We therefore ask whether a motion-pretrained representation transfers to this setting, and quantify how much of the discriminative signal it supplies. Each subject is represented by five task-level motion embeddings, one per clinical task, produced by a frozen pretrained spatio-temporal graph convolutional network (ST-GCN) that fuses the thirteen body-worn sensors into a whole-body embedding; a three-layer Transformer with validity-mask weighting aggregates these tokens for binary PD-versus-control classification on the WearGait-PD cohort (181 subjects: 100 PD, 81 controls). Under a leakage-free nested protocol with repeated subject-disjoint stratified 5-fold cross-validation (5 seeds; 25 estimates per model) and paired significance testing, the model attains a balanced accuracy of 0.834 ± 0.087, macro-F1 of 0.842 ± 0.094, and AUC of 0.842 ± 0.103. It leads six classical baselines and a spectrogram-CNN on accuracy-based metrics, though random forest, gradient boosting, and the spectrogram-CNN edge ahead on AUC; after correction for fold correlation, none of these between-model differences is significant. The one robust finding is a transfer effect: replacing the pretrained encoder with a random one of identical architecture lowers balanced accuracy by 15.5 points when frozen (p = 0.043) and 20.4 when trained end-to-end (p = 0.014). Discrimination is preserved under 1:1 age matching (0.846) and across both genders, so it is not explained by age imbalance. Motion-pretrained skeletal encoders thus supply the majority of the discriminative signal, while the aggregator contributes gains inseparable from noise at this cohort size.
