Related Experiment Video
Updated: Aug 25, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Improving Hybrid CTC/Attention Architecture for Agglutinative Language Speech Recognition
Zeyu Ren1, Nurmemet Yolwas1, Wushour Slamu1
1Xinjiang Multilingual Information Technology Laboratory, Xinjiang Multilingual Information Technology Research Center, College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China.
This study introduces a novel end-to-end (E2E) system for automatic speech recognition (ASR) in agglutinative languages. The proposed method enhances feature extraction and incorporates advanced training techniques, significantly improving performance on Turkish and Uzbek speech datasets.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Speech Processing
Background:
- End-to-end (E2E) automatic speech recognition (ASR) models offer an alternative to traditional systems but require substantial training data.
- Hybrid CTC/attention ASR systems show promise for low-resource conditions but are underutilized for Central Asian languages like Turkish and Uzbek.
Purpose of the Study:
- To enhance the performance of E2E ASR systems for agglutinative languages.
- To address the data scarcity challenge in low-resource ASR scenarios.
Main Methods:
- Dataset augmentation through noise addition and speed perturbation.
- Proposed a novel multi-scale feature extraction method (MSPC) using varied convolution kernel sizes.
- Improved attention mechanisms and utilized CTC objective function with BERT-initialized language models for training and decoding.
Main Results:
- The MSPC feature extractor outperformed VGGnet.
- The proposed model demonstrated accelerated convergence and improved accuracy.
- On LibriSpeech, character error rate (CER) and word error rate (WER) increased by 2.42% and 2.96% respectively.
- Significant WER reductions of 7.07% and 7.08% were achieved on Common Voice-Turkish and Uzbek datasets, respectively.
Conclusions:
- The developed E2E ASR system shows competitive performance, approaching advanced systems.
- The proposed methods are effective in improving ASR for low-resource agglutinative languages.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
12:49Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019