Modeling monophthongal versus diphthongal /aɪ/ in sung vocal performance with interpretable machine learning.

Romeo De Timmerman1, Gil Verbeke1

  • 1Department of Linguistics, Ghent University, Ghent 9000, Belgium.

Summary

This study shows that machine learning can predict if singers produce diphthongal or monophthongal /aɪ/ vowels based on acoustic data. Explainable AI reveals dynamic vocal features differentiate these vowel perceptions.