Related Experiment Videos
Automatic recognition of syllable-final nasals preceded by /epsilon/
P Loizou1, M Dorman, A Spanias
1Department of Electrical Engineering, Arizona State University, Tempe 85287-7206.
The Journal of the Acoustical Society of America
|March 1, 1995
Summary
This study enhances automatic speech recognition for nasal consonants by integrating perceptual findings into hidden Markov models (HMMs). Incorporating vowel-nasal transitions significantly improves the recognition of alveolar sounds.
Area of Science:
- Speech Recognition
- Acoustic Phonetics
- Machine Learning
Background:
- Perceptual studies highlight the importance of nasal murmurs and formant transitions for identifying nasal consonant place of articulation.
- Automatic speech recognition systems often struggle with accurate nasal sound discrimination.
Purpose of the Study:
- To improve automatic recognition of nasal consonants using hidden Markov models (HMMs).
- To integrate acoustic cues identified in perceptual experiments into an HMM-based speech recognition system.
Main Methods:
- Incorporated acoustic segments bordering the nasal release, specifically vowel-nasal transitions, into an HMM-based system.
- Adjusted the HMM recognizer to prioritize vowel-nasal transition segments over nasal murmur and vowel portions for /epsilon m/ and /epsilon n/ syllables.
Main Results:
- Achieved a 7% improvement in alveolar recognition by explicitly modeling vowel-nasal transition segments.
- Realized an additional 6% overall improvement by focusing the HMM on transition segments.
- Attained an average [m]-[n] recognition rate of 95% on an independent test set of 60 speakers.
Conclusions:
- Explicitly modeling vowel-nasal transitions in HMMs significantly enhances nasal consonant recognition.
- Focusing HMMs on perceptually relevant acoustic segments, like transitions, improves overall speech recognition accuracy for nasals.