Related Experiment Video
Updated: Jul 11, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Towards Automated Vocal Mode Classification in Healthy Singing Voice-An XGBoost Decision Tree-Based Machine Learning
Jeroen Sol1, Mathias Aaen2, Cathrine Sadolin3
1Institute for Computing and Information Sciences, Radboud University, Nijmegen, the Netherlands.
Machine learning models can automatically classify singing voice qualities like metallic versus non-metallic sounds. While accurate, these AI systems currently approximate, but do not surpass, human expert judgment in vocal mode discrimination.
Area of Science:
- Acoustics and Audio Signal Processing
- Artificial Intelligence in Music Performance
- Voice Science and Technology
Background:
- Auditory-perceptual assessment of singing voice is standard but suffers from inconsistent reliability.
- Machine learning (ML) offers potential for objective, automated voice analysis.
- Previous ML models show promise in classifying pathological and healthy voices.
Purpose of the Study:
- To develop and evaluate an XGBoost machine learning classifier for automated vocal mode classification in healthy singing voices.
- To compare the performance of the ML model against human auditory-perceptual assessments.
- To identify key acoustic features for accurate singing voice classification.
Main Methods:
- An XGBoost decision tree classifier was trained using various acoustic features: mel-frequency cepstrum coefficients (MFCCs), glottal features, voice quality features, and alpha-ratios.
- The model was tested on distinguishing metallic vs. non-metallic singing and general vocal modes in male and female singers.
- Performance was benchmarked against 41 professional singers assessing 64 vocal samples.
Main Results:
- The ML classifier achieved high accuracy in distinguishing metallic (92% F1-score for males, 87% for females) and vocal modes (70% F1 for males, 69% for females).
- Key features for classification were MFCCs and alpha-ratios, with models using only these showing comparable performance.
- The automated system's performance approximated or was subpar to human expert assessment.
Conclusions:
- XGBoost models show significant potential for automated singing voice analysis, particularly using MFCCs and alpha-ratios.
- Current AI models do not yet match human perceptual discrimination accuracy but represent an improvement over prior automated methods.
- Further research may enhance AI capabilities for reliable and objective singing voice assessment.
Related Concept Videos
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...

