Related Experiment Video
Updated: Aug 5, 2026

04:04
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Explainable Machine Learning for Voice Disorder Screening Using Acoustic Features From Multiple Speech Tasks
Yeon Woo Lee1, Na Ryung Kim2, Beom Yang Shin2
1Department of Speech-Language Pathology and Audiology, Kosin University, Busan, South Korea.
Summary
Machine learning models effectively distinguish pathological voices using combined acoustic features from multiple speech tasks. Explainable AI identified key predictors, supporting its use in voice disorder screening.
Area of Science:
- Speech pathology and acoustics
- Machine learning in healthcare
- Explainable Artificial Intelligence (XAI)
Background:
- Distinguishing pathological voices from normal controls is crucial for early diagnosis and treatment.
- Traditional methods rely on subjective perceptual evaluation, which can be inconsistent.
- Objective acoustic analysis offers a promising avenue for voice disorder assessment.
Purpose of the Study:
- To evaluate machine learning models for pathological voice detection using various acoustic feature sets.
- To compare the performance of different machine learning classifiers.
- To apply Explainable AI techniques to identify key acoustic predictors of dysphonia.
Main Methods:
- Analysis of 1111 voice samples using four acoustic feature sets: sustained vowel, connected speech, concatenated, and combined.
- Training and evaluation of Support Vector Machine, Random Forest, and eXtreme Gradient Boosting classifiers.
- Performance assessment using accuracy and Area Under the ROC Curve, with SHapley Additive exPlanations for feature interpretation.
Main Results:
- The combined feature set with eXtreme Gradient Boosting achieved the highest accuracy (97.31%).
- Acoustic features from multiple speech tasks significantly impacted performance.
- Key predictors included CSID, CPP_v, L/H ratio_v, L/H ratio_SD, and AVQI.
- Misclassifications were limited to mild dysphonia grades.
Conclusions:
- Integrating acoustic data from diverse speech tasks enhances pathological voice discrimination.
- Explainable machine learning successfully identified clinically relevant acoustic predictors.
- This approach shows potential as an objective and transparent tool for voice disorder screening.