Related Experiment Video
Updated: Sep 16, 2025

Functional Near-Infrared Spectroscopy Hyperscanning Study in Psychological Counseling
Published on: January 17, 2025
Research on Acoustic Scene Classification Based on Time-Frequency-Wavelet Fusion Network.
Fengzheng Bi1, Lidong Yang1,2
1School of Digital and Intelligent Industry, Inner Mongolia University of Science and Technology, Baotou 014010, China.
This study introduces a novel time-frequency-wavelet fusion network for acoustic scene classification. The proposed model significantly improves accuracy on urban sound datasets, demonstrating its effectiveness in complex acoustic environments.
Area of Science:
- Artificial Intelligence
- Signal Processing
- Machine Learning
Background:
- Acoustic scene classification identifies environments from sound signals.
- Variations in audio due to different cities and devices challenge model accuracy.
- Existing methods struggle with diverse and complex acoustic environments.
Purpose of the Study:
- To develop an advanced acoustic scene classification model.
- To enhance model robustness against audio variations.
- To improve classification accuracy in real-world scenarios.
Main Methods:
- A time-frequency-wavelet fusion network was proposed.
- A time-frequency-wavelet module extracted features across time, frequency, and wavelet domains.
- Gated temporal-spatial attention and visual state space modules were integrated for enhanced contextual modeling.
- Kolmogorov-Arnold network layers replaced traditional multilayer perceptrons in the classifier.
Main Results:
- Achieved 56.16% average accuracy on the TAU Urban Acoustic Scenes 2022 mobile dataset, a 6.53% improvement over the baseline.
- Reached 97.60% accuracy on the UrbanSound8K dataset, outperforming existing methods.
- Demonstrated significant performance gains in complex acoustic scenarios.
Conclusions:
- The proposed time-frequency-wavelet fusion network effectively improves acoustic scene classification accuracy.
- The model exhibits strong generalization ability across different datasets and complex environments.
- The fusion of multidimensional audio information is crucial for robust acoustic scene recognition.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Discrete Fourier Transform
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Continuous -time Fourier Transform
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...

