Related Experiment Video
Updated: Apr 13, 2026

Computer-based Multitaper Spectrogram Program for Electroencephalographic Data
Published on: November 13, 2019
Separable spectro-temporal Gabor filter bank features: Reducing the complexity of robust features for automatic
Marc René Schädler1, Birger Kollmeier1
1Medizinische Physik and Cluster of Excellence Hearing4all, Universität Oldenburg, D-26111 Oldenburg, Germany.
Separate spectral and temporal processing is sufficient for robust automatic speech recognition (ASR). This finding improves ASR performance by reducing the required signal-to-noise ratio (SNR) and processing time.
Area of Science:
- Speech processing
- Auditory perception
- Machine learning for ASR
Background:
- Robust feature extraction is crucial for automatic speech recognition (ASR) in noisy environments.
- Traditional spectro-temporal Gabor filter banks (GBFB) assume simultaneous processing of spectral and temporal information.
- Investigating the necessity of this simultaneous processing for ASR performance is key.
Purpose of the Study:
- To determine if spectral and temporal processing can be separated for robust ASR feature extraction.
- To introduce and evaluate separate Gabor filter bank (SGBFB) features.
- To compare the performance of SGBFB features against traditional GBFB and Mel-frequency cepstral coefficients (MFCCs).
Main Methods:
- Decomposing a spectro-temporal Gabor filter bank (GBFB) into independent spectral and temporal 1D-Gabor filter banks.
- Extracting separate Gabor filter bank (SGBFB) features using these independent banks.
- Evaluating SGBFB features on the CHiME keywords-in-noise recognition task.
Main Results:
- Spectral and temporal processing can be performed independently for robust ASR.
- SGBFB features allowed a 1.2 dB lower signal-to-noise ratio (SNR) compared to GBFB, improving word error rate by 12.8%.
- The real-time factor of spectro-temporal processing was reduced by over an order of magnitude.
Conclusions:
- Simultaneous spectro-temporal processing is not required for robust ASR feature extraction.
- SGBFB features offer significant improvements in ASR performance and computational efficiency.
- SGBFB features require a lower SNR compared to GBFB and MFCCs for human-listener-equivalent recognition performance.
Related Concept Videos
IR Spectrum Peak Splitting: Symmetric vs Asymmetric Vibrations
Bandpass Sampling
A bandpass signal has a spectrum with a lower frequency limit, denoted as ω1, and an upper frequency limit, denoted as ω2....
Discrete-Time Fourier Series
For a discrete-time periodic signal x[n]...
IR Frequency Region: Fingerprint Region
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Sampling Continuous Time Signal
In the...

