Related Experiment Video
Updated: Nov 12, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Modal and non-modal voice quality classification using acoustic and electroglottographic features.
The COVAREP feature set, particularly harmonic model features, best classifies voice quality. Mel-frequency cepstral coefficients (MFCCs) also showed effectiveness, especially for modal, breathy, and strained voice qualities.
Area of Science:
- Speech and Audio Processing
- Biomedical Engineering
- Machine Learning for Signal Processing
Background:
- Accurate voice quality classification is crucial for diagnosing voice disorders.
- Evaluating different acoustic and physiological voice features is essential for improving classification accuracy.
- Mel-frequency cepstral coefficients (MFCCs) are widely used but their effectiveness across different voice signals needs further investigation.
Purpose of the Study:
- To compare the performance of the COVAREP feature set against MFCCs for voice quality classification.
- To evaluate feature effectiveness across acoustic, glottal inverse filtered (GIF), and electroglottographic (EGG) voice signals.
- To assess the utility of various machine learning classifiers for speaker-independent voice quality analysis.
Main Methods:
- Extracted features including COVAREP (glottal source, frequency warped cepstrum, harmonic model) and MFCCs from voice recordings.
- Utilized acoustic, GIF, and EGG waveforms as input signals.
- Employed Support Vector Machines, Random Forests, Deep Neural Networks, and Gaussian Mixture Models for classification.
- Implemented a leave-one-speaker-out cross-validation strategy for speaker independence.
Main Results:
- The full COVAREP feature set achieved the highest classification accuracy of 79.97%.
- Harmonic model features, a subset of COVAREP, performed strongly with 78.47% accuracy.
- Static+dynamic MFCCs achieved 74.52% accuracy, effectively classifying modal, breathy, and strained qualities from acoustic and GIF signals.
- The EGG waveform showed reduced classification performance compared to other signals.
Conclusions:
- The COVAREP feature set, especially harmonic model features, demonstrates superior performance for voice quality classification compared to MFCCs.
- MFCCs remain a viable option for classifying specific voice quality dimensions from acoustic and GIF signals.
- Speaker-independent classification models can effectively differentiate voice qualities, with variations based on feature sets and signal types.
More Related Videos
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
04:04Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Related Concept Videos
Pulse amplitude and quality
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Physical Assessment of the Respiratory Tract IV: Auscultation
Breath Sounds
Breath sounds are categorized into vesicular, bronchovesicular, and bronchial.
Classification of Skeletal Muscle Fibers
Slow-Twitch Muscle Fibers
Slow oxidative, muscle fibers appear red due to large numbers of capillaries and high levels of...
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...