Related Experiment Video
Updated: Jun 20, 2026

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
Improving Dysarthric Speech Segmentation With Emulated and Synthetic Augmentation
Saeid Alavi Naeini1,2, Leif Simmatis1, Deniz Jafari1,2
1KITE, Toronto Rehabilitation Institute, University Health Network (UHN) Toronto ON M5G 2A2 Canada.
Augmenting speech data with emulated and synthetic variations significantly improves automatic speech recognition for neurological disease diagnosis. This reduces the need for extensive clinical data in speech segmentation tasks.
Area of Science:
- Speech processing
- Neurological disease diagnosis
- Machine learning
Background:
- Acoustic features from speech aid in diagnosing neurological diseases and monitoring symptoms.
- Temporal segmentation of audio signals into words is crucial for feature extraction.
- Automatic Speech Recognition (ASR) and sequence alignment can automate speech segmentation.
Purpose of the Study:
- To explore augmentation methods for improving ASR performance on dysarthric speech.
- To assess the feasibility of using emulated and synthetic data for ASR model fine-tuning.
- To reduce reliance on scarce and privacy-sensitive clinical data.
Main Methods:
- Fine-tuning pre-trained ASR models using two augmentation strategies: 1) healthy speakers altering rate/loudness, 2) synthetic speech with varied rate/accent.
- Evaluating model performance on dysarthric speech data.
- Comparing augmentation methods against fine-tuning with real clinical data.
Main Results:
- Augmentation with emulated and synthetic data outperformed models fine-tuned solely on real clinical data.
- Performance matched models fine-tuned on real data combined with synthetic speech.
- The best model achieved 5.7% word error rate and 94.4% accuracy on neurological disease data.
- Mean intersection-over-union of 89.2% for word segmentation against ground truth.
Conclusions:
- Emulated and synthetic data augmentations effectively reduce the need for real clinical data in ASR model fine-tuning for dysarthric speech.
- These methods enhance speech segmentation accuracy for neurological disease assessment.
- The approach offers a scalable and privacy-preserving solution for clinical speech analysis.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Line Loss
Line loss impacts power delivery efficiency in a balanced three-phase circuit. The symmetry in such a circuit simplifies the...
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...
Reconstruction of Signal using Interpolation
Lossy Lines and Overvoltages
Attenuation
When constant series resistance and shunt conductance are present, voltage and current equations are modified. The propagation constant indicates that voltage and current waves consist of both forward and backward traveling components. These waves attenuate as they propagate, with the attenuation factor related to the resistance and conductance. In a...

