Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Linear Approximation in Frequency Domain01:26

Linear Approximation in Frequency Domain

198
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
198
Linear Approximation in Time Domain01:21

Linear Approximation in Time Domain

171
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
171
Reconstruction of Signal using Interpolation01:10

Reconstruction of Signal using Interpolation

423
Signal processing techniques are essential for accurately converting continuous signals to digital formats and vice versa. When a continuous signal is sampled with a period T, the resulting sampled signal exhibits replicas of the original spectrum in the frequency domain, spaced at intervals equal to the sampling frequency. To handle this sampled signal, a zero-order hold method can be applied, which creates a piecewise constant signal by retaining each sample's value until the next...
423
Sampling Continuous Time Signal01:11

Sampling Continuous Time Signal

438
In signal processing, a continuous-time signal can be sampled using an impulse-train sampling technique, followed by the zero-order hold method. Impulse-train sampling involves the use of a periodic impulse train, which consists of a series of delta functions spaced at regular intervals determined by the sampling period. When a continuous-time signal is multiplied by this impulse train, it generates impulses with amplitudes corresponding to the signal's values at the sampling points.
In the...
438
Classification of Signals01:30

Classification of Signals

1.0K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.0K
Extraction: Advanced Methods00:56

Extraction: Advanced Methods

659
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
659

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Robust frame-level speaker localization guided by multi-channel speech enhancement and inter-channel phase-difference losses.

The Journal of the Acoustical Society of America·2025
Same author

Perceptual effects of reducing algorithmic latency on deep-learning based noise reductiona).

The Journal of the Acoustical Society of America·2025
Same authorSame journal

<math></math> Estimation and Voicing Detection With Cascade Architecture in Noisy Speech.

IEEE/ACM transactions on audio, speech, and language processing·2025
Same author

Leveraging laryngograph data for robust voicing detection in speech.

The Journal of the Acoustical Society of America·2024
Same author

Progress made in the efficacy and viability of deep-learning-based noise reduction.

The Journal of the Acoustical Society of America·2023
Same author

Causal Deep CASA for Monaural Talker-Independent Speaker Separation.

IEEE/ACM transactions on audio, speech, and language processing·2020

Related Experiment Video

Updated: Oct 28, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K

A New Framework for CNN-Based Speech Enhancement in the Time Domain.

Ashutosh Pandey1, DeLiang Wang2

  • 1Department of Computer Science and Engineering, The Ohio State University, Columbus, OH 43210 USA.

IEEE/ACM Transactions on Audio, Speech, and Language Processing
|July 15, 2021
PubMed
Summary

This study introduces a novel learning method for fully convolutional neural networks (CNNs) to enhance speech in the time domain. The approach uses a differentiable frequency domain loss for improved speech enhancement performance.

Keywords:
Speech enhancementdeep learningfully convolutional neural networkmean absolute errortime domain enhancement

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

626
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

533

Related Experiment Videos

Last Updated: Oct 28, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

626
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

533

Area of Science:

  • Signal Processing
  • Machine Learning
  • Artificial Intelligence

Background:

  • Speech enhancement is crucial for improving the intelligibility of noisy audio signals.
  • Traditional methods often struggle with complex noise types and preserving speech quality.
  • Fully convolutional neural networks (CNNs) show promise but require effective training strategies.

Purpose of the Study:

  • To propose a novel learning mechanism for time-domain speech enhancement using CNNs.
  • To leverage frequency-domain analysis within a time-domain CNN training framework.
  • To address limitations of existing speech enhancement techniques.

Main Methods:

  • A fully convolutional neural network (CNN) is employed for time-domain speech enhancement.
  • A differentiable operation converts time-domain signals to the frequency domain during training.
  • Mean absolute error loss is applied to the Short-Time Fourier Transform (STFT) magnitude for training.

Main Results:

  • The proposed method significantly outperforms existing speech enhancement techniques.
  • The CNN operates in the time domain, avoiding the invalid STFT problem.
  • The approach effectively utilizes frequency-domain knowledge for enhanced speech quality.

Conclusions:

  • The novel learning mechanism enables effective time-domain speech enhancement with CNNs.
  • This method offers a robust and implementable solution for speech processing tasks.
  • The approach demonstrates superior performance and applicability to related signal processing challenges.