Related Experiment Video
Updated: Jun 6, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.4K
Enhancing dysarthric speech recognition through SepFormer and hierarchical attention network models with multistage
R Vinotha1, D Hepsiba2, L D Vijay Anand1
1Division of Robotics Engineering, Karunya Institute of Technology and Sciences, Coimbatore, Tamil Nadu, India.
Scientific Reports
|November 28, 2024
Summary
This study enhances dysarthric speech recognition (DSR) by integrating SepFormer-Speech Enhancement Generative Adversarial Network (S-SEGAN) for improved clarity. The combined approach significantly boosts word recognition accuracy, with the Conformer-Hierarchical Attention Network (C-HAN) achieving the highest performance.
Area of Science:
- Speech Technology
- Artificial Intelligence
- Biomedical Engineering
Background:
- Dysarthria significantly impairs speech clarity, posing challenges for Automatic Speech Recognition (ASR) systems.
- Existing ASR systems struggle with the reduced intelligibility characteristic of dysarthric speech.
- Enhancing dysarthric speech recognition (DSR) is crucial for improving communication accessibility for affected individuals.
Purpose of the Study:
- To develop and evaluate a novel system for enhanced dysarthric speech recognition (DSR).
- To integrate advanced speech enhancement techniques as a front-end for DSR systems.
- To assess the impact of different model architectures and data augmentation on DSR accuracy.
Main Methods:
- Integration of SepFormer-Speech Enhancement Generative Adversarial Network (S-SEGAN) for dysarthric speech enhancement (DSE).
- Utilizing a multi-stage transfer learning approach, training on LibriSpeech and fine-tuning on dysarthric speech data.
- Evaluating Transformer and Conformer models, with and without Hierarchical Attention Network (HAN), incorporating DSE and data augmentation.
Main Results:
- Baseline Transformer and Conformer models achieved Word Recognition Accuracy (WRA) of 68.60% and 69.87%.
- Models with Hierarchical Attention Network (HAN) improved WRA to 71.07% (T-HAN) and 73% (C-HAN).
- Integration of DSE and data augmentation led to significant gains, with the Conformer-HAN model achieving a top WRA of 84.07%.
Conclusions:
- The proposed S-SEGAN front-end significantly enhances DSR accuracy when integrated with DSR models.
- Multi-stage transfer learning and model architectures like Conformer-HAN are effective for improving DSR performance.
- The findings demonstrate a substantial advancement in creating more robust and accurate speech recognition systems for individuals with dysarthria.

