Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Amplifying Signals via Enzymatic Cascade01:22

Amplifying Signals via Enzymatic Cascade

When a ligand binds to a cell-surface receptor, the receptor's intracellular domain changes shape, which may either activate its enzyme function or allow its binding to other molecules. The initial signal is amplified by most signal transduction pathways. This means that a single ligand molecule can activate multiple molecules of a downstream target. Proteins that relay a signal are most commonly phosphorylated at one or more sites, activating or inactivating the protein. Kinases catalyze the...
Cascaded Op Amps01:16

Cascaded Op Amps

Operational amplifiers (op-amps) are versatile electronic components that can be interconnected in a cascade - one after another in a linear sequence. This cascading is possible due to their infinite input resistance and zero output resistance, allowing them to maintain their input-output relationships even when connected in series.
In a cascaded system, each op-amp is referred to as a stage. The output of one stage drives the input of the subsequent stage. As the input signal passes through...
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Robust frame-level speaker localization guided by multi-channel speech enhancement and inter-channel phase-difference losses.

The Journal of the Acoustical Society of America·2025
Same author

Perceptual effects of reducing algorithmic latency on deep-learning based noise reductiona).

The Journal of the Acoustical Society of America·2025
Same author

Leveraging laryngograph data for robust voicing detection in speech.

The Journal of the Acoustical Society of America·2024
Same author

Progress made in the efficacy and viability of deep-learning-based noise reduction.

The Journal of the Acoustical Society of America·2023
Same author

A New Framework for CNN-Based Speech Enhancement in the Time Domain.

IEEE/ACM transactions on audio, speech, and language processing·2021
Same author

Causal Deep CASA for Monaural Talker-Independent Speaker Separation.

IEEE/ACM transactions on audio, speech, and language processing·2020

Related Experiment Video

Updated: Jun 19, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Estimation and Voicing Detection With Cascade Architecture in Noisy Speech.

Yixuan Zhang1, Heming Wang1, DeLiang Wang2

  • 1Department of Computer Science and Engineering, Ohio State University, Columbus, OH 43210 USA.

IEEE/ACM Transactions on Audio, Speech, and Language Processing
|May 5, 2025
PubMed
Summary

This study introduces a novel neural cascade architecture for robust pitch tracking in noisy speech. The method jointly optimizes speech enhancement and pitch estimation for improved accuracy, especially in challenging low signal-to-noise ratio conditions.

Keywords:
Complex domain processingdensely-connected convolutional recurrent neural networkmulti-task learningneural cascade architecturepitch trackingvoicing detection

More Related Videos

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)
10:55

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)

Published on: April 12, 2026

Related Experiment Videos

Last Updated: Jun 19, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)
10:55

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)

Published on: April 12, 2026

Area of Science:

  • Speech Processing
  • Acoustics
  • Machine Learning

Background:

  • Pitch tracking is crucial for speech processing but challenging in noisy conditions.
  • Non-stationary noises degrade speech signals, complicating pitch estimation and voicing detection.
  • Accurate voicing detection is essential for reliable pitch tracking.

Purpose of the Study:

  • To propose a neural cascade architecture for joint pitch estimation and voicing detection.
  • To enhance pitch tracking performance in noisy speech.
  • To develop a speaker-independent and noise-independent method.

Main Methods:

  • A neural cascade architecture was designed, integrating speech enhancement and pitch tracking modules.
  • The multi-task framework was trained to optimize both pitch estimation and voicing detection simultaneously.
  • The model was trained in a speaker-independent and noise-independent manner.

Main Results:

  • The integrated enhancement module significantly improved pitch estimation and voicing detection accuracy, particularly at low signal-to-noise ratios (SNRs).
  • The proposed multi-task framework outperformed combined single-task models in performance and efficiency.
  • The method demonstrated robustness across various noise conditions.

Conclusions:

  • The neural cascade architecture offers a superior and efficient approach to pitch tracking in noisy speech.
  • Joint optimization of speech enhancement and pitch tracking is effective for improving accuracy.
  • The proposed method shows significant advantages over existing pitch tracking techniques.