Machine Learning-Based Multiclass Classification of Benign Vocal Fold Lesions Using Acoustic, Clinical, and
Ebru Karakaya Gojayev1, Zahide Çiler Büyükatalay2, Arjin Öksüz3
1University of Health Sciences Türkiye, Ankara Etlik City Hospital, Clinic of Otorhinolaryngology, Ankara, Türkiye.
Objective:
To develop and evaluate a multiclass random forest classifier for distinguishing among five types of benign vocal fold lesions using acoustic and phonatory parameters, patient-reported outcomes (PROs), demographic characteristics, and occupational voice demands.
Methods:
This retrospective study included 217 adults with confirmed diagnoses of vocal fold nodules, vocal fold polyps, intracordal cysts, Reinke's edema, or sulcus vocalis at a tertiary otolaryngology center. Features included five acoustic parameters (fundamental frequency, jitter, shimmer, maximum phonation time, s/z ratio), three PRO instruments, the voice handicap index, voice-related quality of life, and reflux symptom index (RSI), and demographic and occupational variables. A class-weighted random forest classifier was evaluated using stratified five-fold cross-validation with bootstrap-derived 95% confidence intervals; multinomial logistic regression and linear support vector machine served as comparators.
Results:
The random forest classifier outperformed the comparator models, achieving accuracy of 0.576, balanced accuracy of 0.476, macro-F1 of 0.479, and a macro-averaged area under the receiver operating characteristic curve of 0.76. The highest class-level F1 scores were observed for sulcus vocalis (0.703) and vocal fold polyps (0.625); intracordal cysts showed the lowest recall (0.111), reflecting acoustic overlap with vocal fold polyps. Top 2 accuracy was approximately 74%. Fundamental frequency, RSI, and maximum phonation time were the highest-ranked predictors.
Conclusion:
A multidimensional random forest framework incorporating acoustic measures, PROs, and occupational voice demand demonstrated potential for classifying benign vocal fold lesions. Acoustic parameters alone may not adequately distinguish among these lesions, and the inclusion of patient-reported data may improve model performance. A top 2 accuracy of 74% supports utility as a pre-endoscopic decision-support tool.

