Related Experiment Video
Updated: Nov 7, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.8K
Speech Vision: An End-to-End Deep Learning-Based Dysarthric Automatic Speech Recognition System.
Summary
This study introduces Speech Vision (SV), a novel automatic speech recognition (ASR) system for individuals with dysarthria. SV improves speech recognition accuracy by visually analyzing speech, overcoming challenges like phoneme inaccuracies and data scarcity.
Area of Science:
- Computer Science
- Speech Technology
- Assistive Technology
Background:
- Dysarthria impairs speech intelligibility, creating communication barriers and hindering digital device interaction.
- Existing automatic speech recognition (ASR) systems struggle with dysarthric speech due to phoneme variations, limited data, and labeling inaccuracies.
- Effective ASR for dysarthria can significantly enhance communication and digital accessibility for affected individuals.
Purpose of the Study:
- To develop a dysarthric-specific ASR system that overcomes limitations of current technologies.
- To improve the accuracy and usability of ASR for individuals with dysarthria, particularly those with severe impairments.
Main Methods:
- Introduced Speech Vision (SV), a novel ASR system utilizing visual acoustic modeling to analyze speech features.
- Addressed data scarcity using visual data augmentation, synthetic data generation, and transfer learning.
- Benchmarked SV against state-of-the-art dysarthric ASR systems.
Main Results:
- SV demonstrated superior performance, improving recognition accuracy for 67% of speakers in the UA-Speech dataset.
- Significant improvements were observed for individuals with severe dysarthria.
- The visual acoustic modeling approach effectively mitigated phoneme-related challenges.
Conclusions:
- Speech Vision (SV) represents a significant advancement in ASR for dysarthric speech.
- The novel visual approach and data augmentation strategies enhance ASR performance for this population.
- SV has the potential to greatly improve communication and digital interaction for individuals with dysarthria.

