Related Experiment Video
Updated: Aug 10, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Development of Language Models for Continuous Uzbek Speech Recognition System.
Abdinabi Mukhamadiyev1, Mukhriddin Mukhiddinov1, Ilyos Khujayarov2
1Department of Computer Engineering, Gachon University, Sujeong-gu, Seongnam-si 13120, Republic of Korea.
This study introduces UzLM, a new language model for Uzbek, addressing the lack of resources for low-resource languages. UzLM significantly improves Uzbek speech recognition accuracy using neural networks.
Area of Science:
- Natural Language Processing
- Speech Recognition
- Computational Linguistics
Background:
- Large vocabulary automatic speech recognition and natural language processing applications require language models.
- Research on pre-trained language models predominantly focuses on high-resource languages, neglecting low-resource languages like Uzbek.
- There is a scarcity of publicly available Uzbek speech datasets, hindering the development of language models for this language.
Purpose of the Study:
- To develop a low-resource language model for the Uzbek language.
- To address the limitations in Uzbek natural language processing by creating a robust Uzbek language model.
- To investigate and understand linguistic occurrences specific to the Uzbek language.
Main Methods:
- Proposed UzLM, an Uzbek language model, by evaluating statistical and neural-network-based approaches.
- Developed an Uzbek-specific linguistic representation to enhance the robustness of the language model.
- Utilized a corpus of 80 million words, comprising approximately 68,000 unique words and 15 million sentences, for training.
Main Results:
- The developed UzLM leverages Uzbek-specific linguistic features for improved performance.
- The model was trained using a substantial corpus, demonstrating efficient use of training data compared to previous studies.
- Experimental results on continuous Uzbek speech recognition showed a reduction in character error rate to 5.26% when using neural-network-based models compared to manual encoding.
Conclusions:
- The development of UzLM represents a significant advancement for Uzbek natural language processing and speech recognition.
- The study highlights the effectiveness of neural-network-based language models for low-resource languages.
- UzLM provides a foundation for future research and applications in Uzbek speech technology.
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Components of Language
Language and Cognition
Basic Continuous Time Signals
The unit step function, denoted u(t), is zero for negative time values and one for positive time values, exhibiting a discontinuity at t=0. This function often represents abrupt changes, such as the step voltage introduced when turning a car's...
Sampling Continuous Time Signal
In the...
Multi-input and Multi-variable systems
In the absence...

