Related Experiment Video
Updated: Oct 17, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.7K
Optimising Speaker-Dependent Feature Extraction Parameters to Improve Automatic Speech Recognition Performance for
Marco Marini1, Nicola Vanello1, Luca Fanucci1
1Department of Information Engineering, University of Pisa, Via G. Caruso 16, 56122 Pisa, Italy.
Sensors (Basel, Switzerland)
|October 13, 2021
Summary
A novel speech analysis technique enhances automatic speech recognition (ASR) for individuals with dysarthria by optimizing spectral analysis windows. This method shows promise in improving ASR accuracy for impaired speech, with potential correlations found between voice features and optimal parameters.
Area of Science:
- Speech processing
- Computational linguistics
- Biomedical engineering
Background:
- Automatic Speech Recognition (ASR) systems struggle with dysarthric speech.
- Standard ASR approaches are often ineffective for impaired speech.
- Dysarthria presents significant challenges in human-computer interaction.
Purpose of the Study:
- To validate a new speech analysis technique for improving ASR performance in speakers with dysarthria.
- To investigate correlations between speaker voice characteristics and optimal ASR parameters.
- To enhance the effectiveness of speaker-dependent ASR systems for impaired speech.
Main Methods:
- Fine-tuning spectral analysis window size and shift parameters for Short-Time Fourier Transform (STFT).
- Utilizing a speaker-dependent ASR system.
- Experimenting with Italian speech data from 30 dysarthric speakers (IDEA database) and 10 unimpaired speakers (CLIPS database).
Main Results:
- The proposed speech analysis technique significantly improves ASR performance for speakers with moderate to severe dysarthria.
- The new approach is ineffective for unimpaired or mildly impaired speech.
- A correlation was identified between specific speaker voice features and the optimal window/shift parameters for minimizing ASR errors.
Conclusions:
- The novel speech analysis method offers a viable solution for enhancing ASR in the context of dysarthria.
- Personalized parameter optimization based on voice features can further improve ASR accuracy for dysarthric speakers.
- This research contributes to more inclusive and effective speech technology for individuals with speech impairments.

