A deep learning approach for acoustic-based identification of muscle tension dysphonia and spasmodic dysphonia
Zhou Zhou1,2, Yuan Cheng3,4,5, Qingyi Ren1
1Department of Otolaryngology Head Neck Surgery, Guangdong Provincial People's Hospital, Guangdong Academy of Medical Sciences, Southern Medical University, Guangzhou, China.
Problem:
Differentiating between spasmodic dysphonia (SD), a neurological disorder, and muscle tension dysphonia (MTD), a behavioral voice disorder, based on auditory perception alone is a common clinical challenge. This diagnostic difficulty can lead to delays in appropriate treatment, highlighting the need for objective and reliable assistive diagnostic tools.
Aim:
This study aims to develop and validate an artificial intelligence (AI) model based on deep learning to automatically differentiate between healthy voices, SD, and MTD using only voice audio recordings, and to compare its diagnostic performance against human experts.
Methods:
A retrospective analysis was conducted on 1,597 voice samples (595 healthy, 471 MTD, 531 SD). Voice audio was processed into Log-Mel spectrograms. Pre-trained convolutional neural networks (CNNs), including VGG16, ResNet50, and DenseNet161, were employed for transfer learning to perform both binary (Healthy vs. Disordered) and ternary (Healthy vs. MTD vs. SD) classification. The model's performance was evaluated on a held-out test set and compared to the diagnostic assessments of four otolaryngologists.
Results:
The AI model achieved an accuracy of 89.5% (AUC = 0.956) in distinguishing healthy voices from disordered ones. For the more complex ternary classification, the model attained an accuracy of 71.6%, with class-specific AUCs of 0.957 (Healthy), 0.731 (MTD), and 0.855 (SD). This performance surpassed that of human experts, who achieved average accuracies of 78.2% in binary classification and 60.6% in ternary classification on the same test set.
Conclusion:
The deep learning model trained on a Mandarin pathological voice dataset achieves favorable performance in distinguishing SD from MTD using only voice audio signals, with classification results comparable to those of experienced clinical specialists. This technology serves as a promising objective auxiliary tool for the preliminary screening and differential diagnosis of voice disorders.
