Related Experiment Video
Updated: Sep 10, 2025

08:05
Design and Analysis for Fall Detection System Simplification
Published on: April 6, 2020
10.8K
Deep Spectrogram Learning for Gunshot Classification: A Comparative Study of CNN Architectures and Time-Frequency
Pafan Doungpaisan1, Peerapol Khunarsa2
1Faculty of Industrial Technology and Management, King Mongkut's University of Technology North Bangkok, Bangkok 10800, Thailand.
Journal of Imaging
|August 27, 2025
Summary
Deep learning models accurately classify firearm sounds using spectrogram images. Mel, CQT, and Cochleagram spectrograms with CNNs achieved over 94% accuracy for public safety applications.
Area of Science:
- Acoustics
- Machine Learning
- Signal Processing
Background:
- Gunshot sound classification is vital for public safety, forensic analysis, and surveillance.
- Deep learning models offer potential for improving the accuracy and efficiency of firearm sound identification.
Purpose of the Study:
- To evaluate deep learning model performance for classifying firearm sounds.
- To investigate the effectiveness of various time-frequency spectrogram representations converted into images for CNN analysis.
Main Methods:
- Analyzed twelve time-frequency spectrograms (Mel, Bark, MFCC, CQT, Cochleagram, STFT, FFT, Reassigned, Chroma, Spectral Contrast, Wavelet).
- Converted spectrograms into RGB images for application of computer vision techniques.
- Trained six Convolutional Neural Network (CNN) architectures (ResNet18, ResNet50, ResNet101, GoogLeNet, Inception-v3, InceptionResNetV2) on spectrogram images.
Main Results:
- CQT, Cochleagram, and Mel spectrograms achieved high classification accuracy (>94%) when used with deep CNNs like ResNet101 and InceptionResNetV2.
- Transforming spectrograms into images enabled effective use of image-based processing and deep learning models.
- The approach demonstrated robustness in capturing spectral-temporal patterns for accurate firearm sound classification.
Conclusions:
- Deep learning models, particularly CNNs, combined with image-based spectrogram analysis, provide a powerful framework for firearm sound classification.
- Specific spectrogram representations (CQT, Cochleagram, Mel) are highly effective when converted to image formats for deep learning.
- This methodology enhances accuracy and robustness in applications like public safety and forensic investigations.
Related Concept Videos
Classification of Signals
878
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
878
Force Classification
1.6K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.6K
Discrete Fourier Transform
404
The Discrete Fourier Transform (DFT) is a fundamental tool in signal processing, extending the discrete-time Fourier transform by evaluating discrete signals at uniformly spaced frequency intervals. This transformation converts a finite sequence of time-domain samples into frequency components, each representing complex sinusoids ordered by frequency. The DFT translates these sequences into the frequency domain, effectively indicating the magnitude and phase of each frequency component present...
404
IR Frequency Region: Fingerprint Region
1.1K
IR spectra are divided into two main regions: the diagnostic region and the fingerprint region. The diagnostic region of the spectrum lies above 1500 cm−1. The absorptions resulting from single-bond vibrations of the N–H, C–H, and O–H stretch at higher wavenumbers and appear on the left side of the spectrum. The stretching absorptions of the C≡C and C≡N occur between 2100–2300 cm−1. In contrast, those arising from stretching absorptions of the...
1.1K
Difference from Background: Limit of Detection
7.1K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
7.1K

