Related Experiment Video
Updated: Oct 28, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
A New Framework for CNN-Based Speech Enhancement in the Time Domain
Ashutosh Pandey1, DeLiang Wang2
1Department of Computer Science and Engineering, The Ohio State University, Columbus, OH 43210 USA.
This study introduces a novel learning method for fully convolutional neural networks (CNNs) to enhance speech in the time domain. The approach uses a differentiable frequency domain loss for improved speech enhancement performance.
Area of Science:
- Signal Processing
- Machine Learning
- Artificial Intelligence
Background:
- Speech enhancement is crucial for improving the intelligibility of noisy audio signals.
- Traditional methods often struggle with complex noise types and preserving speech quality.
- Fully convolutional neural networks (CNNs) show promise but require effective training strategies.
Purpose of the Study:
- To propose a novel learning mechanism for time-domain speech enhancement using CNNs.
- To leverage frequency-domain analysis within a time-domain CNN training framework.
- To address limitations of existing speech enhancement techniques.
Main Methods:
- A fully convolutional neural network (CNN) is employed for time-domain speech enhancement.
- A differentiable operation converts time-domain signals to the frequency domain during training.
- Mean absolute error loss is applied to the Short-Time Fourier Transform (STFT) magnitude for training.
Main Results:
- The proposed method significantly outperforms existing speech enhancement techniques.
- The CNN operates in the time domain, avoiding the invalid STFT problem.
- The approach effectively utilizes frequency-domain knowledge for enhanced speech quality.
Conclusions:
- The novel learning mechanism enables effective time-domain speech enhancement with CNNs.
- This method offers a robust and implementable solution for speech processing tasks.
- The approach demonstrates superior performance and applicability to related signal processing challenges.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:04Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
Related Concept Videos
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
Reconstruction of Signal using Interpolation
Sampling Continuous Time Signal
In the...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Extraction: Advanced Methods