Related Experiment Videos
Spontaneous speech recognition using a statistical coarticulatory model for the vocal-tract-resonance dynamics
1Department of Electrical and Computer Engineering, University of Waterloo, Ontario, Canada. deng@microsoft.com
The Journal of the Acoustical Society of America
|January 6, 2001
Summary
A new statistical coarticulatory model enhances spontaneous speech recognition by incorporating vocal tract dynamics. This model outperforms traditional Hidden Markov Models (HMMs) by efficiently capturing long-range speech context.
Area of Science:
- Speech Recognition
- Acoustic Phonetics
- Statistical Modeling
Background:
- Conventional Hidden Markov Models (HMMs) struggle with long-span context dependence in spontaneous speech.
- Incorporating dynamic vocal tract behavior is crucial for accurate speech recognition.
Purpose of the Study:
- To develop a novel statistical coarticulatory model for spontaneous speech recognition.
- To leverage dynamic vocal tract resonance for improved model design and performance.
Main Methods:
- Formulated a constrained, nonstationary, nonlinear dynamic system for speech acoustics.
- Developed a generalized Expectation-Maximization (EM) algorithm for parameter learning.
- Utilized spontaneous speech data from the Switchboard corpus for experiments.
Main Results:
- The new coarticulatory model demonstrated consistently superior performance compared to a benchmark HMM system.
- Experiments validated the model's effectiveness in spontaneous speech recognition and synthesis.
- Analysis revealed the model's strength in capturing target-directed vocal tract behavior and long-span context.
Conclusions:
- The proposed statistical coarticulatory model offers a more parsimonious and effective approach to spontaneous speech recognition.
- The model's inherent structure successfully represents complex speech dynamics and context.
- This advancement holds promise for improving the accuracy and robustness of speech recognition technologies.