Automatic speech recognition for Telugu: a comparative analysis of Wav2Vec 2.0 model variants and hyperparameter

Anvita Manne1, Nikhita James1, Ishaan Jain1

  • 1Department of Quantum AI, School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, India.

Summary

Developing automatic speech recognition (ASR) for Telugu, a low-resource language, requires robust models. Fine-tuning Wav2Vec 2.0 variants showed XLS-R-1B achieved the best performance, highlighting model size and hyperparameter importance for effective Telugu ASR.