Related Experiment Video
Updated: Feb 26, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Automatic assessment of voice quality using machine learning
Yat Chun Au1, Nan Yan2, Manwa L Ng1
1Speech Science Laboratory, Faculty of Education, University of Hong Kong, Hong Kong, China.
Objectives:
This study aimed to develop and validate machine learning (ML) models for automated prediction of perceptual dysphonia severity, as indexed by the Grade (G) parameter of the GRBAS scale, using acoustic analyses of sustained vowels. The overarching goal was to enhance objectivity, reproducibility, and efficiency in clinical voice assessment.
Methods:
A total of 524 sustained/a/samples were collected from three databases. The evaluations of all recordings by ten raters using the GRBAS scale, that achieved excellent interrater reliability (Krippendorff's α = 0.96), were modelled. Forty-seven acoustic features spanning spectral, cepstral, perturbation, and noise-based indices were extracted using Parselmouth (Praat). Five ML classifiers-Decision Tree (DT), Random Forest (RF), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), and Categorical Boosting (CatBoost)-were trained using 5-fold cross-validation (80/20 split) and evaluated by accuracy, F1-score, and quadratic weighted kappa (QWK).
Results:
Gradient boosting algorithms outperformed traditional tree-based models. LightGBM achieved the highest QWK (0.945), followed by CatBoost (QWK = 0.941) and XGBoost (QWK = 0.935). Feature-importance analyses identified cepstral measures-particularly Smoothed Cepstral Peak Prominence (CPPS), Cepstral Spectral Index of Dysphonia (CSID), Acoustic Voice Quality Index (AVQI), Harmonics-to-Noise Ratio (HNR) as the most influential predictors of perceptual Grade (G), while jitter and shimmer parameters contributed minimally. Correlation analyses confirmed strong associations between Grade and AVQI (r = 0.854), HNR (r = -0.853), and cepstral indices (r = -0.835 to -0.832).
Conclusions:
Gradient boosting methods, particularly LightGBM, produced near-expert agreement with perceptual ratings, supporting their potential as objective, interpretable tools for clinical dysphonia assessment.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
05:48Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Pulse amplitude and quality
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...