Related Experiment Video
Updated: Aug 6, 2026

Dynamic Digital Biomarkers of Motor and Cognitive Function in Parkinson's Disease
Published on: July 24, 2019
Longitudinal voice biomarker trajectory modelling for Parkinson's disease severity: domain-adaptive transfer learning
Deepika Roselind Johnson1, G Logeswari1
1School of Computer Science and Engineering, Vellore Institute of Technology, Chennai Campus, Chennai, India.
Introduction:
Speech and voice changes affect up to 90% of people with Parkinson's disease (PD), a progressive neurodegenerative disorder affecting approximately 10 million people worldwide. Although continuous monitoring of disease severity is clinically important, most existing voice-based computational approaches focus on binary PD-versus-control classification and do not model longitudinal symptom progression. To address this gap, we propose a domain-adaptive transformer model, DAT-PD, for predicting continuous PD severity trajectories from real-world smartphone voice recordings.
Methods:
DAT-PD was developed using the public mPower dataset, comprising 58,247 voice recordings from 5,800 participants. The proposed pipeline included noise-aware acoustic preprocessing, extraction of extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) features, a domain-adaptive attention mechanism to reduce cross-device and cross-environment variability and a longitudinal trajectory decoder. The model was trained, inferred, and evaluated using continuous MDS-UPDRS Part-II scores as the sole prediction target. Confounder-aware domain adaptation was incorporated to address demographic imbalance, including the age gap between the PD cohort and healthy controls. Robustness was further evaluated under harsh acoustic conditions with signal-to-noise ratios as low as 0 dB.
Results:
On the held-out test set, DAT-PD achieved a mean absolute error (MAE) of 2.74 MDS-UPDRS units (95% CI: 2.44-3.01), root mean squared error (RMSE) of 3.61 (95% CI: 3.18-4.04), and R² of 0.93 (95% CI: 0.91-0.95), outperforming six state-of-the-art baseline models. eGeMAPS features substantially outperformed MFCC-only representations, reducing MAE from 5.21 to 2.74. SHAP-based explainability identified MFCC-2, Shimmer (APQ5) and Jitter as the most influential longitudinal voice biomarkers.
Discussion:
The superior performance of DAT-PD suggests that domain-adaptive longitudinal modeling can effectively capture clinically meaningful voice-based severity trajectories in PD. The advantage of eGeMAPS over MFCC-only features is likely due to its ability to represent phonatory and prosodic characteristics relevant to PD dysarthria, including F0 dynamics, loudness contour, shimmer, jitter and spectral flux. By maintaining robustness under noisy real-world acoustic conditions, DAT-PD supports unsupervised home-based monitoring using standard smartphones. These findings align with the Bridge2AI-Voice research agenda and position DAT-PD as a clinically implementable, non-invasive tool for continuous PD severity assessment.
