Related Experiment Video
Updated: Jul 16, 2026

Manufacturing Process for Non-Adhesive Super-Soft Vocal Fold Models
Published on: January 5, 2024
Modeling monophthongal versus diphthongal /aɪ/ in sung vocal performance with interpretable machine learning
Romeo De Timmerman1, Gil Verbeke1
1Department of Linguistics, Ghent University, Ghent 9000, Belgium.
Abstract:
This study combines modern source-separation techniques with interpretable machine learning methods to investigate how binary perceptual annotations of diphthongal and monophthongal /aɪ/ relate to measurable acoustic variation in sung vocal performance. Using a corpus of studio-recorded music processed with vocal-instrumental source separation, we extract F1 and F2 trajectories for 1004 /aɪ/ tokens, each perceptually categorized as either monophthongal or diphthongal. These trajectories were modeled using two approaches: (i) a gradient boosted decision tree trained on engineered acoustic features and (ii) a multilayer perceptron (MLP) trained directly on raw formant trajectories. Both models achieved high accuracy and Area Under the Receiver Operating Characteristic Curve, indicating that perceptual labels can reliably be predicted from both raw and engineered acoustic input. Moreover, a Shapley Additive Explanations-based feature importance analysis showed that features such as ΔF1/ΔF2, cubic spline coefficients, and trajectory derivatives captured systematic differences between perceived monophthongal and diphthongal tokens, highlighting the value of dynamic representations over static distance-based measures. The results indicate that perceptual annotations of diphthongal and monophthongal /aɪ/ correspond with quantifiable acoustic information and demonstrate how explainable machine learning can help map gradient vowel dynamics onto binary perceptual categories. The study further emphasizes how recent advances in source separation make sung performance a valuable new domain for phonetic research.
