Related Experiment Video
Updated: Nov 21, 2025

08:45
Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
14.9K
A convolutional recurrent neural network with attention framework for speech separation in monaural recordings.
Chao Sun1, Min Zhang2, Ruijuan Wu3
1College of Electrical Engineering, Sichuan University, Chengdu, 610065, China.
Scientific Reports
|January 15, 2021
Summary
This study introduces a novel Convolutional Recurrent Neural Network with Attention (CRNN-A) for monaural speech separation. The CRNN-A framework significantly improves speech separation quality by combining CNN and RNN strengths with an attention mechanism.
Area of Science:
- Signal Processing
- Artificial Intelligence
- Machine Learning
Background:
- Monaural speech separation often yields unsatisfactory results due to limitations of single network architectures.
- Existing methods struggle to effectively learn features for high-quality speech separation.
Purpose of the Study:
- To propose a novel Convolutional Recurrent Neural Network with Attention (CRNN-A) framework for enhanced monaural speech separation.
- To fuse the complementary strengths of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) for improved feature learning.
Main Methods:
- A CNN front-end with dual convolution kernels extracts time- and frequency-domain features from spectrograms.
- Features are concatenated and processed through convolutional layers before being fed to an RNN back-end.
- An attention mechanism is integrated to focus on relationships between feature maps, optimizing separation performance.
Main Results:
- The CRNN-A framework demonstrated superior performance compared to baseline RNN and other popular speech separation methods.
- Evaluations on the MIR-1K dataset showed significant improvements in Global Normalised Source-to-Distortion Ratio (GNSDR), Global Source-to-Interference Ratio (GSIR), and Global Source-to-Artifacts Ratio (GSAR).
Conclusions:
- The proposed CRNN-A framework effectively integrates CNN and RNN capabilities, enhanced by an attention mechanism, for superior speech separation.
- This framework offers a promising approach for speech separation, speech enhancement, and related audio processing fields.
