Related Experiment Video
Updated: Jul 16, 2026

06:24
Manufacturing Process for Non-Adhesive Super-Soft Vocal Fold Models
Published on: January 5, 2024
Modeling monophthongal versus diphthongal /aɪ/ in sung vocal performance with interpretable machine learning.
Romeo De Timmerman1, Gil Verbeke1
1Department of Linguistics, Ghent University, Ghent 9000, Belgium.
The Journal of the Acoustical Society of America
|July 15, 2026
Summary
This study shows that machine learning can predict if singers produce diphthongal or monophthongal /aɪ/ vowels based on acoustic data. Explainable AI reveals dynamic vocal features differentiate these vowel perceptions.
Area of Science:
- Phonetics
- Acoustic Phonetics
- Computational Linguistics
Background:
- Investigating vowel perception and production in singing is crucial for understanding vocal performance.
- Diphthongal and monophthongal vowel distinctions are perceptually significant but acoustically complex.
- Source-separation and machine learning offer new tools for analyzing sung vocalizations.
Purpose of the Study:
- To determine the relationship between acoustic variations and perceptual categories of the /aɪ/ vowel in sung performances.
- To apply interpretable machine learning to analyze formant trajectories and their correlation with vowel perception.
- To explore the utility of source separation techniques in phonetic research on sung vocalizations.
Main Methods:
- Utilized source separation to isolate vocal tracks from studio-recorded music.
- Extracted F1 and F2 formant trajectories for 1004 /aɪ/ tokens categorized perceptually.
- Employed gradient boosted decision trees and multilayer perceptrons (MLPs) to model formant data.
- Applied Shapley Additive Explanations (SHAP) for feature importance analysis.
Main Results:
- Both machine learning models achieved high accuracy in predicting perceptual vowel categories.
- Acoustic features like trajectory derivatives and spline coefficients significantly differentiated monophthongal and diphthongal /aɪ/.
- Dynamic acoustic features proved more informative than static measures for distinguishing vowel types.
- Perceptual labels for /aɪ/ vowels reliably correspond to quantifiable acoustic information.
Conclusions:
- Perceptual judgments of monophthongal and diphthongal /aɪ/ in singing are supported by measurable acoustic differences.
- Interpretable machine learning effectively maps continuous acoustic variations to discrete perceptual categories.
- Sung vocal performance is a promising new area for phonetic research, enhanced by advanced audio processing techniques.
