Related Experiment Videos
Intonation and dialog context as constraints for speech recognition.
1Center for Speech Technology Research, University of Edinburgh, U.K. pault@cstr.ed.ac.uk
Language and Speech
|April 4, 2000
Summary
This study enhances automatic speech recognition (ASR) by using intonation and dialog context. Integrating move-specific language models and dialog act recognition significantly reduces word error rates in task-oriented speech.
Area of Science:
- Computational Linguistics
- Speech Processing
- Artificial Intelligence
Background:
- Automatic Speech Recognition (ASR) systems typically use bigram language models.
- Task-oriented dialog speech presents unique challenges for ASR performance.
- Dialog context and intonation are often underutilized in conventional ASR systems.
Purpose of the Study:
- To improve automatic speech recognition (ASR) performance by incorporating intonation and dialog context.
- To develop and evaluate move-specific language models for task-oriented dialog.
- To investigate the impact of automatic dialog move type recognition on ASR word error rates.
Main Methods:
- Utilized the DCIEM Maptask corpus of spontaneous task-oriented dialog speech.
- Developed separate bigram language models for each of the 12 dialog move types.
- Combined an intonation model, a dialog model, and speech recognizer likelihoods for move type determination.
Main Results:
- Move-specific language models significantly reduced word error rate when the correct model was applied.
- The integrated system, combining automatic move type recognition with move-specific models, reduced overall word error rate compared to a baseline.
- Word error rate improvements were primarily observed in 'initiating' move types, while 'response' move types showed good recognition of the type itself.
Conclusions:
- Integrating intonation and dialog context, through move-specific language models and dialog act recognition, offers a significant improvement for ASR in task-oriented dialog.
- The approach effectively enhances word recognition for critical dialog segments.
- Further research can explore more sophisticated dialog and intonation modeling for broader ASR applications.