Differentiating Etiologies of Dysphonia: A Machine Learning Approach With Multidomain Acoustic Features
1Speech Science Laboratory, Faculty of Education, University of Hong Kong, Hong Kong, China.
Objectives:
The present study aims to develop and evaluate the feasibility of automatic categorical differentiation of pathological voice disorders (organic, functional, and neurologic) using machine learning (ML) models based on multidomain acoustic features from connected speech.
Methods:
A total of 584 pathological recordings obtained from the Advanced Voice Function Assessment Databases were analyzed. Twenty-nine prespecified acoustic features covering multiple domains including prosodic, perturbation, harmonicity, and cepstral-spectral domains, were extracted. Three classifiers, including Random Forest, Extreme Gradient Boosting (XGBoost), and CatBoost, plus a soft-voting ensemble, were trained with stratified 5-fold cross-validation on an 80% development split and evaluated on a 20% external hold-out set. Performance metrics including accuracy, precision (positive predictive value [PPV]), recall/sensitivity, specificity, F1, per-class indices, and confusion matrices, were assessed. Feature importance (XGBoost) was summarized via Weight, Gain, and Cover to support physiologic interpretability.
Results:
Results indicated that XGBoost yielded the best overall performance (accuracy = 0.82, PPV = 0.80, sensitivity = 0.70, specificity = 0.91, F1 = 0.73), followed by CatBoost (accuracy = 0.79, F1 = 0.66) and the ensemble (accuracy = 0.79, F1 = 0.66); random forest showed lower sensitivity (0.55) and F1 (0.60). Across models, organic dysphonia achieved the highest F1 (0.86-0.87), neurologic was moderate (F1 ≈ 0.70-0.72), and functional was most challenging (F1 = 0.42-0.60). Feature importance was dominated by cepstral-spectral measures, particularly cepstral peak prominence smoothed, cepstral spectral index of dysphonia, and harmonics-to-noise ratio, indicating that indices related to signal periodicity and glottal closure-related structure help were influential in the etiologic classification task.
Conclusion:
Gradient-boosting ML, XGBoost in particular, enables clinically meaningful multiclass differentiation of voice disorder etiologies from sentence-level acoustics. The prominence of cepstral-spectral predictors provides physiologically plausible markers for objective assessment. ML-assisted acoustics can complement perceptual and laryngoscopic evaluation for triage and longitudinal monitoring; improving sensitivity for functional dysphonia will likely require task diversity, multimodal features, and larger, multilingual validation cohorts.

