Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Linear Approximation in Frequency Domain01:26

Linear Approximation in Frequency Domain

79
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
79
Determination of Expected Frequency01:08

Determination of Expected Frequency

2.1K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.1K
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

160
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
160
Sampling Continuous Time Signal01:11

Sampling Continuous Time Signal

176
In signal processing, a continuous-time signal can be sampled using an impulse-train sampling technique, followed by the zero-order hold method. Impulse-train sampling involves the use of a periodic impulse train, which consists of a series of delta functions spaced at regular intervals determined by the sampling period. When a continuous-time signal is multiplied by this impulse train, it generates impulses with amplitudes corresponding to the signal's values at the sampling points.
In the...
176
Classification of Signals01:30

Classification of Signals

314
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
314
Continuous -time Fourier Transform01:11

Continuous -time Fourier Transform

237
The Fourier series is instrumental in representing periodic functions, offering a powerful method to decompose such functions into a sum of sinusoids. This technique, however, necessitates modification when applied to nonperiodic functions. Consider a pulse-train waveform consisting of a series of rectangular pulses. When these pulses have a finite period, they can be accurately represented by a Fourier series. Yet, as the period approaches infinity, resulting in a single, isolated pulse, the...
237

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Co-removal of norfloxacin and Cr(VI) by Co/N-doped carbon activating peroxymonosulfate: <sup>1</sup>O<sub>2</sub> oxidation coupled with interfacial electron transfer.

Environmental research·2026
Same author

A speech prediction model based on codec modeling and transformer decoding.

Computer speech & language·2026
Same author

A Molecular Trimming Strategy for Hypoxia-Tolerant Photosensitizers With Enhanced cGAS-STING Activation.

Angewandte Chemie (International ed. in English)·2026
Same author

Towards decoupling frontend enhancement and backend recognition in monaural robust ASR.

Computer speech & language·2026
Same author

Colocalization of eQTLs With Type 2 Diabetes and Glycemic Traits Using Whole-Genome Sequences in Diverse Populations From the NHLBI Trans-Omics in Precision Medicine (TOPMed) Program.

Diabetes·2026
Same author

Embodied Transboundary Greenhouse Gas Emissions and Mitigation Opportunities of over 700 Belt and Road Initiative Projects.

Environmental science & technology·2026

Related Experiment Video

Updated: May 9, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.3K

Densely-connected Convolutional Recurrent Network for Fundamental Frequency Estimation in Noisy Speech.

Yixuan Zhang1, Heming Wang1, DeLiang Wang1,2

  • 1Department of Computer Science and Engineering, The Ohio State University, USA.

Interspeech
|May 2, 2025
PubMed
Summary

This study introduces a new deep learning model for fundamental frequency (F0) estimation in noisy speech. The proposed cascade model improves F0 estimation accuracy, especially in challenging low signal-to-noise ratio conditions.

Keywords:
Pitch trackingcascade architecturecomplex domaindensely-connected convolutional recurrent neural network

More Related Videos

Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology
09:44

Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology

Published on: March 8, 2024

4.5K
fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
11:15

fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals

Published on: May 23, 2017

7.2K

Related Experiment Videos

Last Updated: May 9, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.3K
Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology
09:44

Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology

Published on: March 8, 2024

4.5K
fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
11:15

fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals

Published on: May 23, 2017

7.2K

Area of Science:

  • Speech processing
  • Machine learning
  • Signal analysis

Background:

  • Accurate fundamental frequency (F0) estimation is crucial for speech synthesis and analysis.
  • Estimating F0 from noisy speech remains challenging due to noise corrupting the speech's harmonic structure.

Purpose of the Study:

  • To develop a robust F0 estimation method for noisy speech using deep learning.
  • To investigate the impact of speech enhancement on F0 estimation accuracy.
  • To propose a cascade model for improved F0 estimation in adverse acoustic conditions.

Main Methods:

  • Framing F0 estimation as a multi-class classification problem.
  • Utilizing a frequency-domain densely-connected convolutional neural network (DC-CRN).
  • Comparing complex Short-Time Fourier Transform (STFT) with magnitude STFT as input features.
  • Developing a cascade model integrating speech enhancement and F0 estimation modules.

Main Results:

  • The DC-CRN model significantly outperforms baseline methods in F0 detection rate.
  • Complex STFT input yields better performance than magnitude STFT.
  • A cascade model integrating speech enhancement further improves F0 estimation accuracy.
  • The cascade model shows particular benefits in low signal-to-noise ratio (SNR) environments.

Conclusions:

  • Deep learning, specifically DC-CRN, offers a powerful approach for F0 estimation in noisy speech.
  • Speech enhancement, when integrated effectively in a cascade model, can enhance F0 estimation performance.
  • The proposed cascade model provides a promising solution for robust F0 estimation under various noise conditions.