Related Experiment Video
Updated: May 11, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
The role of binary mask patterns in automatic speech recognition in background noise
1Department of Computer Science and Engineering and Center for Cognitive Science, The Ohio State University, Columbus, Ohio 43210, USA. narayaar@cse.ohio-state.edu
Ideal binary masks significantly enhance automatic speech recognition (ASR) in noisy conditions, even at very low signal-to-noise ratios (SNRs). This study reveals ASR performance closely mirrors human speech recognition capabilities.
Area of Science:
- Speech processing
- Acoustic signal analysis
- Machine learning for audio
Background:
- Automatic Speech Recognition (ASR) performance degrades significantly in noisy environments.
- Ideal binary masks (IBMs) are known to improve ASR by separating target speech from noise.
- The specific impact of IBM patterns on ASR across varying noise conditions and vocabulary sizes remains under-explored.
Purpose of the Study:
- To investigate the role of binary mask patterns in ASR performance.
- To analyze the influence of noise, signal-to-noise ratios (SNRs), and vocabulary size on ASR with binary masking.
- To compare ASR performance with binary masking to human speech intelligibility.
Main Methods:
- Binary masks were computed using two criteria: local criterion (LC) based on SNR, and comparison of local target energy to long-term average spectral energy.
- ASR experiments were conducted under various noise conditions, SNRs, and vocabulary sizes.
- ASR performance metrics were analyzed and compared to human intelligibility data.
Main Results:
- Binary masking substantially improved ASR accuracy, even at extremely low SNRs (-60 dB).
- ASR performance patterns under different conditions qualitatively matched human speech intelligibility.
- The difference between LC and mixture SNR was a stronger predictor of ASR accuracy than LC alone.
- Optimal performance was achieved at an LC below 0 dB, contrary to maximizing SNR gain.
Conclusions:
- Binary masking is a highly effective technique for improving ASR in challenging noisy environments.
- ASR systems employing binary masks exhibit performance characteristics surprisingly similar to human listeners.
- Maximizing SNR gain is not necessarily the optimal strategy for enhancing either human or machine speech recognition in noise.
More Related Videos
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Automatic Processing and Automatic Social Behavior
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

