Related Experiment Video
Updated: Sep 21, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.6K
Automatic Speech Recognition Method Based on Deep Learning Approaches for Uzbek Language.
Abdinabi Mukhamadiyev1, Ilyos Khujayarov2, Oybek Djuraev3
1Department of Computer Engineering, Gachon University, Sujeong-gu, Seongnam-si 13120, Korea.
Sensors (Basel, Switzerland)
|May 28, 2022
Summary
This study introduces novel speech recognition models for the Uzbek language, achieving a 14.3% word error rate. The research addresses the gap in low-resource language processing for speech technologies.
Area of Science:
- Natural Language Processing
- Artificial Intelligence
- Speech Technology
Background:
- Speech recognition systems predominantly focus on high-resource languages like English, neglecting low-resource languages.
- Uzbek language and its dialects lack comprehensive speech recognition models, limiting technological accessibility.
- Effective communication and globalization necessitate advancements in speech recognition for diverse linguistic backgrounds.
Purpose of the Study:
- To develop and evaluate advanced speech recognition models for the Uzbek language.
- To address the underrepresentation of low-resource languages in current speech recognition research.
- To improve the accuracy and efficiency of speech recognition systems for Uzbek.
Main Methods:
- Implementation of an End-To-End Deep Neural Network-Hidden Markov Model (DNN-HMM) for Uzbek speech recognition.
- Development of a hybrid Connectionist Temporal Classification (CTC)-attention network to enhance model training.
- Utilizing the CTC objective function within the attention model to optimize training time and accuracy.
Main Results:
- The proposed hybrid CTC-attention model achieved a word error rate (WER) of 14.3% on the Uzbek language dataset.
- The model demonstrated improved speech recognition accuracy and reduced training time compared to existing methods.
- Performance was evaluated using a newly collected Uzbek language dataset comprising 207 hours of recordings.
Conclusions:
- The developed speech recognition models show significant promise for the Uzbek language and its dialects.
- The hybrid CTC-attention approach offers an effective strategy for improving low-resource speech recognition systems.
- This research contributes to bridging the gap in speech technology accessibility for underrepresented languages.

