Related Experiment Videos
Evaluation of formant-like features on an automatic vowel classification task
Febe de Wet1, Katrin Weber, Louis Boves
1Department of Language and Speech, University of Nijmegen, Nijmegen, The Netherlands. F.de.wet@let.kun.nl
The Journal of the Acoustical Society of America
|October 14, 2004
Summary
This study compared automatically extracted formant-like speech features to hand-labeled formants for vowel classification. While formant-like features performed well in clean, gender-dependent conditions, they were outperformed by Mel-frequency cepstral coefficients (MFCCs) in noisy and gender-independent scenarios.
Area of Science:
- Speech processing
- Acoustic phonetics
- Machine learning for speech recognition
Background:
- Automatic speech recognition (ASR) seeks effective low-dimensional speech signal representations.
- The comparison of automatically extracted formant-like features to true formants is crucial for ASR development.
- Existing research lacks direct comparisons of robust formants and HMM2 features against hand-labeled formants.
Purpose of the Study:
- To compare the performance of two automatically extracted formant-like features (robust formants and HMM2 features) against hand-labeled formants.
- To evaluate these features in a vowel classification task under clean and noisy conditions.
- To benchmark against Mel-frequency cepstral coefficients (MFCCs) as a state-of-the-art ASR feature.
Main Methods:
- Utilized a subset of the American English vowels database with hand-labeled formants.
- Extracted robust formant features using the split Levinson algorithm.
- Extracted HMM2 features via two-dimensional hidden Markov models for speech signal frequency segmentation.
- Included Mel-frequency cepstral coefficients (MFCCs) for comparison.
Main Results:
- Formant-like features showed comparable performance to hand-labeled formants in clean, gender-dependent vowel classification.
- Performance of formant-like features was inferior to hand-labeled formants in gender-independent classification.
- Formant-like features demonstrated a lack of inherent noise robustness in acoustic conditions.
- MFCCs achieved comparable or superior results across all conditions, despite higher dimensionality.
Conclusions:
- Automatically extracted formant-like features show potential but have limitations in gender-independent and noisy conditions.
- MFCCs offer a more robust and effective, albeit higher-dimensional, alternative for ASR.
- Further research is needed to improve the noise robustness and generalization of formant-based speech features.