Related Experiment Video
Updated: Aug 5, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Explainable Machine Learning for Voice Disorder Screening Using Acoustic Features From Multiple Speech Tasks
Yeon Woo Lee1, Na Ryung Kim2, Beom Yang Shin2
1Department of Speech-Language Pathology and Audiology, Kosin University, Busan, South Korea.
Objective:
This study evaluated and compared the performance of machine learning models across four acoustic feature sets designed to distinguish pathological voices from normal controls. Additionally, Explainable Artificial Intelligence techniques were applied to identify the specific acoustic parameters driving the models' predictions.
Methods:
A total of 1111 voice samples were analyzed. Four feature sets were constructed: sustained vowel, connected speech, concatenated sample, and the combined feature set. Support Vector Machine, Random Forest, and eXtreme Gradient Boosting classifiers were trained and evaluated using an independent test set. Model performance was assessed via accuracy and Area Under the ROC Curve, with SHapley Additive exPlanations utilized for post hoc feature interpretation.
Results:
The combined feature set achieved the highest classification performance, with eXtreme Gradient Boosting yielding the best overall results (accuracy = 97.31%). Differences in feature sets exerted a greater impact on performance than the choice of classifier, highlighting the critical role of acoustic information acquired from multiple speech tasks. CSID, CPP_v, L/H ratio_v, L/H ratio_SD, and AVQI emerged as the most influential predictors. Notably, all misclassifications occurred only between perceptually normal (grade 0) and mildly dysphonic (grade 1) voices, whereas no misclassifications were observed in moderate (grade 2) or severe (grade 3) dysphonia.
Conclusion:
Integrating acoustic information from multiple speech tasks significantly improves the discrimination of pathological voices. The deployment of explainable machine learning identified a compact, clinically meaningful subset of acoustic predictors, strongly supporting its potential as an objective and transparent screening tool for voice disorders.