Related Experiment Videos
Multiexpert automatic speech recognition using acoustic and myoelectric signals
Adrian D C Chan1, Kevin B Englehart, Bernard Hudgins
1Department of Systems and Computer Engineering, Carleton University, Ottawa, ON K1S 5B6, Canada. adcchan@sce.carleton.ca
IEEE Transactions on Bio-Medical Engineering
|April 11, 2006
Summary
This study introduces a novel multiexpert automatic speech recognition (ASR) system combining acoustic and facial myoelectric signals (MES). The plausibility method significantly enhances ASR robustness in noisy environments.
Area of Science:
- Speech Recognition
- Signal Processing
- Machine Learning
Background:
- Conventional automatic speech recognition (ASR) systems exhibit reduced accuracy in noisy acoustic conditions.
- Robustness in ASR is crucial for reliable performance across diverse environments.
- Integrating multimodal information can potentially improve ASR system performance.
Purpose of the Study:
- To develop and evaluate a multiexpert ASR system that combines acoustic and facial myoelectric signals (MES).
- To introduce and assess the 'plausibility method' for expert combination within an ASR framework.
- To compare the plausibility method against Borda count and score-based methods for combining ASR experts.
Main Methods:
- A multiexpert ASR system was implemented, integrating acoustic speech data with facial MES data.
- The plausibility method, based on evidence theory, was used to combine acoustic and MES ASR experts.
- Data from 5 subjects with a 10-word vocabulary were collected across an 18-dB noise range.
Main Results:
- The acoustic-only ASR expert's accuracy decreased to 11.5% in high noise.
- The multiexpert system using the plausibility method maintained accuracies above 78.8% across noise levels.
- The plausibility method outperformed individual experts and other combination methods, except at 9-dB noise.
Conclusions:
- The plausibility method effectively enhances ASR robustness by integrating acoustic and MES data.
- Multimodal ASR systems, particularly those employing the plausibility method, offer significant improvements in noisy conditions.
- Facial myoelectric signals provide valuable complementary information for robust speech recognition.