Related Experiment Videos
Automatic Prediction of Vocal Strain Scores in Singing Voice Using Audio and Electroglottographic Modalities
Yuanyuan Liu1, Okko Räsänen2, Tero Ikävalko1
1Speech and Voice Research Laboratory, Tampere University, Finland.
Purpose:
This study developed machine learning models to predict perceptual strain scores in the singing voice using audio and electroglottographic (EGG) recordings. The study examined the predictive capability of distinct feature sets extracted from audio and EGG modalities and assessed the contributions of participant metadata (META) and feature selection to model performance.
Method:
Data were split into mutually exclusive train-validation (n = 16 singers, n ≈ 450 samples) and independent test (n = 11 singers, n ≈ 240 samples) sets to ensure singer-independent generalization. Extracted features included mel-frequency cepstral coefficients, extended Geneva minimalistic acoustic parameter set, wavelet scattering coefficients, and domain-specific descriptors (audio: amplitude modulation; EGG: glottal dynamics), along with singer META. Using leave-one-singer-out cross-validation, we trained regression models (support vector regressor, random forest, and ridge) to predict the expert strain ratings. Recursive feature elimination was employed to optimize feature subsets, and feature-level fusion was implemented to assess multimodal integration.
Results:
During the training-validation phase, the highest Spearman correlations were ρ = .823 (p < .001) for the audio-only model and ρ = .777 (p < .001) for the EGG-only model. On the held-out test set, the audio-only model achieved ρ = .812 (p < .001), while the EGG-only model reached ρ = .759 (p < .001). The optimal overall performance on the test data set (ρ = .825, p < .001) was achieved by a ridge model integrating selected features from audio and META.
Conclusions:
Despite the higher numerical performance of audio-based models, EGG features demonstrated robust predictive potential and superior interpretability within the context of professional singing. The results confirm that combining multimodal data with rigorous feature selection provides a robust framework for the objective assessment of singing voice strain. Direct linkage analysis further verified the tight coupling between glottal and acoustic parameters, grounding the multimodal approach in proven vocal physics.