EnViTSA: Ensemble of Vision Transformer with SpecAugment for Acoustic Event Classification

Kian Ming Lim1, Chin Poo Lee1, Zhi Yang Lee2

  • 1Faculty of Information Science and Technology, Multimedia University, Melaka 75450, Malaysia.

PubMed
Summary

EnViTSA, using Vision Transformers and SpecAugment, improves Acoustic Event Classification by reducing overfitting and enhancing accuracy on benchmark datasets.