Related Experiment Video
Updated: Jul 10, 2026

04:04
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Automatic speech recognition for Telugu: a comparative analysis of Wav2Vec 2.0 model variants and hyperparameter
Anvita Manne1, Nikhita James1, Ishaan Jain1
1Department of Quantum AI, School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, India.
Frontiers in Artificial Intelligence
|July 9, 2026
Summary
Developing automatic speech recognition (ASR) for Telugu, a low-resource language, requires robust models. Fine-tuning Wav2Vec 2.0 variants showed XLS-R-1B achieved the best performance, highlighting model size and hyperparameter importance for effective Telugu ASR.
Area of Science:
- Computational Linguistics
- Speech Processing
- Artificial Intelligence
Background:
- Automatic Speech Recognition (ASR) systems are crucial for digital accessibility and interaction.
- Telugu, despite its speaker base, lacks standardized, noise-robust ASR resources, hindering development.
- Existing ASR models often perform poorly on low-resource languages like Telugu due to optimization for high-resource languages.
Purpose of the Study:
- To fine-tune pre-trained Wav2Vec 2.0 models for Telugu ASR.
- To evaluate the performance of different model variants (XLS-R 300M, XLSR-53, XLS-R-1B) on a Telugu speech corpus.
- To investigate the impact of hyperparameter tuning on ASR accuracy for low-resource languages.
Main Methods:
- Fine-tuning three Wav2Vec 2.0 variants (XLS-R 300M, XLSR-53, XLS-R-1B) on a 48.4-hour Telugu speech corpus.
- Utilizing a curated dataset from Mozilla Common Voice, OpenSLR, and Hugging Face.
- Conducting an ablation study on four hyperparameters and evaluating performance using Word Error Rate (WER) and Character Error Rate (CER).
Main Results:
- The fine-tuned XLS-R-1B model achieved the best performance with a WER of 36.23% and CER of 15.44%.
- XLSR-53 and XLS-R-300M also showed competitive results, with WERs of 37.78% and 37.87% respectively.
- Hyperparameter selection and model size were identified as significant factors influencing ASR performance in Telugu.
Conclusions:
- The XLS-R-1B model demonstrates superior effectiveness for Telugu ASR among the tested Wav2Vec 2.0 variants.
- Optimizing hyperparameters and considering model complexity are critical for developing accurate ASR systems for low-resource languages.
- This research contributes valuable insights and resources for advancing Telugu ASR technology.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...