Related Experiment Video
Updated: Nov 12, 2025

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
Learning Complex Spectral Mapping with Gated Convolutional Recurrent Networks for Monaural Speech Enhancement.
1Department of Computer Science and Engineering, The Ohio State University, Columbus, OH, 43210-1277 USA.
This study introduces a gated convolutional recurrent network (GCRN) for speech enhancement. The GCRN effectively estimates speech phase, significantly improving intelligibility and quality over existing methods.
Area of Science:
- Speech processing
- Machine learning
- Signal processing
Background:
- Speech phase is crucial for perceptual quality but difficult to estimate directly.
- Existing methods like magnitude spectral mapping and complex ratio masking have limitations.
Purpose of the Study:
- To propose a novel gated convolutional recurrent network (GCRN) for complex spectral mapping in monaural speech enhancement.
- To improve both speech intelligibility and perceptual quality by simultaneously enhancing magnitude and phase.
Main Methods:
- Developed a causal gated convolutional recurrent network (GCRN) inspired by multi-task learning.
- Applied complex spectral mapping to estimate real and imaginary spectrograms of clean speech from noisy speech.
Main Results:
- The proposed GCRN significantly outperformed a convolutional neural network (CNN) for complex spectral mapping.
- Achieved higher STOI (Speech Transmission Quality Index) and PESQ (Perceptual Evaluation of Speech Quality) scores compared to other methods.
- Demonstrated effective phase estimation capabilities.
Conclusions:
- Complex spectral mapping with the proposed GCRN is a highly effective approach for monaural speech enhancement.
- The GCRN architecture offers substantial improvements in speech intelligibility and quality.
- This method provides a viable solution for accurate speech phase estimation.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Sampling Continuous Time Signal
In the...
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
Chunking and Rehearsal in Sensory Memory
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....

