Related Experiment Video
Updated: Aug 23, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
End-to-End Deep Convolutional Recurrent Models for Noise Robust Waveform Speech Enhancement.
Rizwan Ullah1, Lunchakorn Wuttisittikulkij1, Sushank Chaudhary1
1Wireless Communication Ecosystem Research Unit, Department of Electrical Engineering, Chulalongkorn University, Bangkok 10330, Thailand.
This study introduces efficient deep learning models for speech enhancement, significantly improving audio quality and clarity. These compact models offer better performance with fewer resources, making them ideal for real-time applications.
Area of Science:
- Speech processing
- Deep learning
- Signal processing
Background:
- End-to-end deep learning (E2E-DL) models are popular for speech enhancement due to their simple design.
- While effective, creating resource-efficient and compact models for real-time processing remains a challenge.
- Modeling sequential and local characteristics of speech signals is crucial for improving E2E model performance.
Purpose of the Study:
- To present resource-efficient and compact neural models for end-to-end, noise-robust, waveform-based speech enhancement.
- To integrate Convolutional Encode-Decoder (CED) and Recurrent Neural Networks (RNNs) within the Convolutional Recurrent Network (CRN) framework for diverse speech enhancement systems.
Main Methods:
- Developed novel Convolutional Recurrent Network (CRN) models combining Convolutional Encode-Decoder (CED) and Recurrent Neural Networks (RNNs).
- Trained and tested models using various noise types, speakers, and datasets (LibriSpeech, DEMAND).
- Evaluated model performance based on quality, intelligibility, trainable parameters, model complexity, and inference time.
Main Results:
- Proposed models achieved improved speech quality (31.61%) and intelligibility (17.18%) compared to noisy speech.
- Demonstrated reduced model complexity and inference time compared to existing recurrent and convolutional models.
- Cross-corpus analysis confirmed the generalization capability of the proposed end-to-end speech enhancement (E2E SE) models.
Conclusions:
- The developed CRN-based E2E SE models offer a superior balance of performance and efficiency.
- These models provide significant improvements in speech quality and intelligibility with reduced computational demands.
- The findings highlight the effectiveness of integrating CED and RNNs for noise-robust speech enhancement.
Related Concept Videos
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Amplifying Signals via Enzymatic Cascade
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Reconstruction of Signal using Interpolation
Sampling Continuous Time Signal
In the...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...

