Related Experiment Video
Updated: Jun 12, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Towards decoupling frontend enhancement and backend recognition in monaural robust ASR
Yufeng Yang1, Ashutosh Pandey1, DeLiang Wang1,2
1Department of Computer Science and Engineering, The Ohio State University, 2015 Neil Avenue, Columbus, 43210, OH, United States.
New speech enhancement models improve automatic speech recognition in noisy conditions. These systems decouple enhancement and recognition, outperforming models trained directly on noisy data for robust ASR.
Area of Science:
- Speech processing
- Machine learning
- Signal processing
Background:
- Speech enhancement (SE) algorithms can improve noisy speech intelligibility.
- Monaural SE has historically underperformed compared to direct noisy speech training for automatic speech recognition (ASR).
- A gap persists between SE advancements and robust ASR system development.
Purpose of the Study:
- To bridge the gap between SE and ASR by proposing novel SE models.
- To enable ASR systems trained solely on clean speech to perform effectively in noisy environments.
- To advance robust ASR through improved SE frontends.
Main Methods:
- Developed three SE models: Attentive Recurrent Network (ARN) in the time-domain, TF-CrossNet in the time-frequency domain, and MP-SENet based on magnitude-phase.
- Decoupled SE frontend from the ASR backend, training the ASR only on clean speech.
- Evaluated performance on WSJ, CHiME-2, LibriSpeech, and CHiME-4 corpora.
Main Results:
- ARN, TF-CrossNet, and MP-SENet significantly improved ASR performance in noisy and reverberant conditions.
- The proposed systems outperformed baseline ASR models trained directly on corrupted speech.
- Achieved state-of-the-art results on CHiME-2 (5.6% WER) and CHiME-4 (3.3/4.4% WER), generalizing to real acoustic scenarios.
Conclusions:
- The proposed SE models effectively eliminate the divide between SE and ASR.
- Decoupled SE frontends allow for robust ASR systems trained on clean speech.
- These advancements pave the way for more effective and generalizable robust ASR.
Related Concept Videos
Auditory Perception
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Downstream Processing
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking the...
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
