Related Experiment Video
Updated: Sep 27, 2026

Dynamic Digital Biomarkers of Motor and Cognitive Function in Parkinson's Disease
Published on: July 24, 2019
A Pilot Study of VGGish-CNN as a Model for the Classification of Parkinson's Disease from a Control Group Using
Mehdi Rashidi1, Syed Adil Hussain Shah2,3, Marco Greco4
1Department of Mathematics and Physics "E. De Giorgi", University of Salento, Via Lecce-Arnesano, 73100 Lecce, Italy.
Abstract:
Introduction: Voice-based digital biomarkers have emerged as a promising, non-invasive approach for the early detection and monitoring of neurodegenerative disorders, particularly Parkinson's disease (PD). Although voice recordings can be acquired easily using mobile health technologies, their integration into routine clinical practice remains limited due to challenges related to model interpretability and clinical validation. This study aimed to investigate the diagnostic potential of voice recordings acquired through the TALIA smartphone and web-based platform and to improve model transparency using explainable artificial intelligence (XAI) techniques. Methods: Voice recordings from participants with PD and control group (CG) were acquired over a six-month period using the TALIA digital-health platform at Vito Fazzi Hospital in Lecce, Italy. The data were evaluated cross-sectionally at the recording level. Sustained vowel (/a/) phonations were preprocessed and transformed into spectrogram images. A transfer-learning framework based on the pre-trained VGGish convolutional neural network (VGGish-CNN) was developed to classify PD and CG voice recordings. To enhance interpretability, explainable artificial intelligence (XAI) methods, including Local Interpretable Model-Agnostic Explanations (LIME) and Occlusion Sensitivity, were integrated to identify the spectro-temporal regions contributing most strongly to model predictions. Results: The proposed VGGish-CNN framework demonstrated excellent classification performance in distinguishing PD from CG. The model achieved an accuracy of 0.93, precision of 0.91, recall of 0.95, F1-score of 0.93, loss of 0.16, and an area under the receiver operating characteristic curve (AUC) of 0.987. XAI analyses provided qualitative, sample-specific visualizations of the spectro-temporal regions contributing to individual model predictions, thereby improving the transparency and interpretability of the deep-learning model. Conclusions: The findings demonstrate that a transfer-learning approach based on VGGish-CNN could achieve promising recording-level classification performance in distinguishing voice recordings from participants with PD and CG within this pilot dataset. Furthermore, XAI techniques provide qualitative insight into the model's decision-making process. These preliminary findings support further investigation of voice-based approaches for PD screening in larger cohorts using participant-level and external validation.
More Related Videos
06:19Integration of Animal Behavioral Assessment and Convolutional Neural Network to Study Wasabi-Alcohol Taste-Smell Interaction
Published on: August 16, 2024
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025