Related Experiment Videos
Speaker normalization of static and dynamic vowel spectral features
1Department of Electrical and Computer Engineering, Old Dominion University, Norfolk, Virginia 23508-0369.
The Journal of the Acoustical Society of America
|July 1, 1991
Summary
Speaker normalization techniques significantly improve automatic vowel classification. Both linear transformation and polynomial warping enhance spectral features, leading to higher accuracy rates for formants and DCTCs.
Area of Science:
- Speech processing
- Acoustic phonetics
- Machine learning for speech
Background:
- Speaker variability in acoustic features poses challenges for automatic speech recognition.
- Vowel spectral features, including formants and DCTCs, are crucial for distinguishing vowels.
- Existing methods often struggle with speaker-independent vowel identification.
Purpose of the Study:
- To evaluate two speaker normalization methods for vowel spectral features.
- To assess the impact of normalization on automatic vowel classification accuracy.
- To compare the effectiveness of different spectral parameters and feature types (static vs. dynamic).
Main Methods:
- Speaker normalization using multivariable linear transformation.
- Speaker normalization using polynomial frequency warping.
- Evaluation with formants and DCTCs (Discrete Cosine Transform Coefficients) as spectral parameters.
- Testing static and dynamic features, with and without fundamental frequency (F0).
Main Results:
- Both normalization methods consistently increased automatic vowel classification rates across all conditions.
- Linear transformation slightly outperformed polynomial warping.
- Normalized DCTCs trajectories achieved the highest classification rate (91%).
- Adding F0 improved non-normalized features but yielded diminishing returns with normalized features.
Conclusions:
- Speaker normalization is effective in improving automatic vowel recognition.
- Linear transformation is a slightly more effective normalization technique.
- Dynamic features combined with normalization show the most promise for robust vowel classification.