Related Experiment Video
Updated: Nov 7, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Deep neural network-based generalized sidelobe canceller for dual-channel far-field speech recognition
Guanjun Li1, Shan Liang2, Shuai Nie2
1National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, China.
A novel deep neural network generalized sidelobe canceller (GSC) optimizes automatic speech recognition (ASR) performance by jointly training with acoustic models. This approach significantly improves speech enhancement and noise robustness in far-field scenarios.
Area of Science:
- Speech processing
- Machine learning
- Acoustic signal processing
Background:
- Traditional generalized sidelobe cancellers (GSC) optimize signal levels, not automatic speech recognition (ASR) performance.
- Far-field ASR systems require robust speech enhancement to mitigate noise.
- Existing GSC methods do not guarantee optimal ASR outcomes.
Purpose of the Study:
- To propose a novel deep neural network (DNN)-based GSC (nnGSC) optimized for ASR performance.
- To develop a GSC structure that integrates learnable modules and joint optimization with acoustic models.
- To enhance noise robustness and performance of far-field ASR systems.
Main Methods:
- Developed a dual-channel DNN-based GSC (nnGSC) structure.
- Initialized nnGSC with traditional GSC coefficients for guided network learning.
- Employed joint optimization between the GSC and the acoustic model.
- Enabled frame-by-frame target direction-of-arrival (DOA) tracking without external algorithms.
- Trained nnGSC with multi-geometry data to improve robustness against array mismatches.
Main Results:
- nnGSC achieved significant relative character error rate (CER) improvements: 23.7% over microphone observation, 13.5% over oracle direction-based super-directive beamformer, 12.2% over oracle direction-based traditional GSC, and 5.9% over oracle mask-based MVDR beamformer.
- The proposed nnGSC automatically tracks target DOA without additional localization.
- nnGSC demonstrated improved robustness against array geometry mismatches.
Conclusions:
- The novel nnGSC structure optimizes ASR performance by integrating DNNs and joint acoustic model training.
- nnGSC offers superior speech enhancement and noise robustness for far-field ASR compared to traditional methods.
- The approach leverages both signal processing knowledge and data-driven learning for effective speech recognition enhancement.
More Related Videos
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Lateralization
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...

