Related Experiment Video
Updated: Jul 10, 2025

05:16
Flying Insect Detection and Classification with Inexpensive Sensors
Published on: October 15, 2014
25.2K
EnViTSA: Ensemble of Vision Transformer with SpecAugment for Acoustic Event Classification
Kian Ming Lim1, Chin Poo Lee1, Zhi Yang Lee2
1Faculty of Information Science and Technology, Multimedia University, Melaka 75450, Malaysia.
Sensors (Basel, Switzerland)
|November 25, 2023
Summary
EnViTSA, using Vision Transformers and SpecAugment, improves Acoustic Event Classification by reducing overfitting and enhancing accuracy on benchmark datasets.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Signal Processing
Background:
- Deep learning models for Acoustic Event Classification (AEC) face overfitting challenges due to high complexity.
- Existing methods struggle with data scarcity and generalization in acoustic event detection.
Purpose of the Study:
- To introduce EnViTSA, an innovative approach for Acoustic Event Classification.
- To enhance AEC performance by mitigating overfitting and addressing data scarcity.
Main Methods:
- Ensemble of pre-trained Vision Transformers combined with SpecAugment data augmentation.
- Transformation of raw acoustic signals into Log Mel-spectrograms.
- Application of time and frequency masking via SpecAugment to generate synthetic training data.
Main Results:
- Achieved 93.50% accuracy on ESC-10, 85.85% on ESC-50, and 83.20% on UrbanSound8K.
- Demonstrated significant reduction in overfitting through the ensemble approach.
- Validated the effectiveness of SpecAugment for AEC data augmentation.
Conclusions:
- EnViTSA offers a substantial advancement in Acoustic Event Classification.
- Vision Transformers and SpecAugment show strong potential for acoustic domain applications.
- The proposed method effectively tackles key challenges in deep learning for AEC.
Related Concept Videos
Types Of Transformers
983
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
983
Perception of Sound Waves
4.5K
The human ear is not equally sensitive to all frequencies in the audible range. It may perceive sound waves with the same pressure but different frequencies as having different loudness. Moreover, the perception of sound waves depends on the health of an individual's ears, which decays with age. The health of one's ears may also be affected by regular exposure to loud noises.
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
4.5K

