Related Experiment Video
Updated: May 22, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
474
Two-stage data augmentation for improved ASR performance for dysarthric speech
Chitralekha Bhat1, Helmer Strik2
1Centre for Language and Speech Technology (CLST), Radboud University Nijmegen, The Netherlands.
Computers in Biology and Medicine
|March 14, 2025
Summary
This study introduces a novel two-stage data augmentation method to improve automatic speech recognition (ASR) for dysarthric speech. The technique significantly reduces word error rate (WER) by leveraging unique speech characteristics.
Area of Science:
- Speech Technology
- Artificial Intelligence
- Biomedical Engineering
Background:
- Accurate automatic speech recognition (ASR) for dysarthric speech is challenging due to limited data.
- Machine learning (ML) and Deep Neural Networks (DNN) show promise but require substantial data.
- Dysarthric speech presents unique characteristics that hinder standard ASR model training.
Purpose of the Study:
- To develop and evaluate a novel two-stage data augmentation scheme for improving ASR performance on dysarthric speech.
- To address the data scarcity issue in training ML/DNN models for dysarthric speech recognition.
- To enhance the accuracy of speaker-dependent ASR systems for individuals with dysarthria.
Main Methods:
- Implemented a two-stage data augmentation: static augmentation (perturbations, devoicing, voice conversion) and dynamic augmentation (Dysarthric SpecAugment).
- Utilized a Transformer acoustic model for an end-to-end ASR system.
- Pre-trained an acoustic model using augmented healthy speech and then fine-tuned it with dysarthric speech data.
Main Results:
- Achieved an absolute improvement of 10.7% in word error rate (WER).
- Demonstrated a relative improvement of 29.2% in WER compared to a baseline without augmentation.
- Attained a final speaker-dependent WER of 25.9% on the UA dysarthric speech corpus.
Conclusions:
- The proposed two-stage data augmentation scheme effectively enhances ASR performance for dysarthric speech.
- Leveraging dysarthric speech characteristics in augmentation is crucial for improving recognition accuracy.
- This approach offers a viable solution to the data scarcity problem in dysarthric ASR.

