Related Experiment Video
Updated: Jan 17, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Convolution-Augmented Transformers for Enhanced Speaker-Independent Dysarthric Speech Recognition
Abstract:
Dysarthria is a motor speech disorder characterized by muscle movement difficulties that complicate verbal communication. It poses significant challenges to Automatic Speech Recognition (ASR) systems due to data scarcity and speaker variability among dysarthric individuals. This study investigates speaker-independent (SI) approaches to assist speakers with communication impairments. Firstly, we developed dysarthric SI models using a Conformer-based system and a three-stage transfer-learning pipeline that employs a selective layer freezing PEFT strategy to mitigate data scarcity. We pre-trained on standard speech and progressively adapted the models to two dysarthric datasets, respectively. Secondly, we introduced a benchmark framework for evaluating the generalizability of SI models with cross-dataset validation-a previously unexplored approach in dysarthric ASR, providing a more realistic scenario. The results demonstrate that the proposed dysarthric SI models outperform all baseline systems. Specifically, on the TORGO dataset, our models improved word recognition accuracy by 21.9% for isolated speech and reduced the word error rate by 18.5% for continuous speech. On UA-Speech, our optimal dysarthric SI model achieved a word recognition improvement of 14.6% over Whisper and 28.3% over the base model for isolated speech. Nevertheless, our cross-dataset testing showed that models tended to produce isolated words when asked to transcribe continuous speech for severe dysarthria, highlighting the need to further improve SI generalization.
Related Concept Videos
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
Convolution: Math, Graphics, and Discrete Signals
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
Convolution Properties I
The commutative property reveals that the input and the impulse response of an LTI (Linear Time-Invariant) system can be interchanged without affecting the output:
Reconstruction of Signal using Interpolation
The Ideal Transformer
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's tangential...
