Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Auditory Perception01:17

Auditory Perception

The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the cochlea, a...
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Downstream Processing01:29

Downstream Processing

Downstream processing begins once fermentation is complete and involves a series of steps to recover and purify products such as acids, vitamins, antibiotics, or proteins.Cell HarvestingFor example, for intracellular protein-based products, the first step is harvesting the cells. This is typically achieved using centrifugation or filtration to separate the cells from the liquid phase.Cell Disruption for Intracellular ProductsIf the target product is intracellular, the harvested cells must be...
Auditory Pathway01:15

Auditory Pathway

Auditory pathways constitute the complex neural circuits responsible for transmitting and interpreting auditory information from the peripheral auditory system to the brain. Sound waves are initially captured by the outer ear, funneled through the ear canal, and reach the tympanic membrane (eardrum). These vibrations are transmitted via the middle ear's ossicles to the inner ear's cochlea.
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking the...
Downsampling01:20

Downsampling

When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same authorSame journal

A speech prediction model based on codec modeling and transformer decoding.

Computer speech & language·2026
Same author

Explainable ensemble learning framework for predicting industrial and energy sector GHG emissions.

Scientific reports·2026
Same author

A Molecular Trimming Strategy for Hypoxia-Tolerant Photosensitizers With Enhanced cGAS-STING Activation.

Angewandte Chemie (International ed. in English)·2026
Same author

Polyphenols-rich Indian barberry berries extract alleviates inorganic arsenic exposure-induced cognitive impairments and associated gut microflora alterations.

Food research international (Ottawa, Ont.)·2026
Same author

Granitic intrusions enhance strain localization and rapid mantle exhumation along an oceanic detachment fault.

Science advances·2026
Same author

Efficacy of SWIM technology combined with direct aspiration first pass technique for large vessel occlusion in acute ischemic stroke.

American journal of translational research·2026

Related Experiment Video

Updated: Jun 12, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Towards decoupling frontend enhancement and backend recognition in monaural robust ASR.

Yufeng Yang1, Ashutosh Pandey1, DeLiang Wang1,2

  • 1Department of Computer Science and Engineering, The Ohio State University, 2015 Neil Avenue, Columbus, 43210, OH, United States.

Computer Speech & Language
|June 11, 2026
PubMed
Summary

New speech enhancement models improve automatic speech recognition in noisy conditions. These systems decouple enhancement and recognition, outperforming models trained directly on noisy data for robust ASR.

Keywords:
CHiME-2CHiME-4Robust ASRSpeech distortionSpeech enhancement

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

Related Experiment Videos

Last Updated: Jun 12, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

Area of Science:

  • Speech processing
  • Machine learning
  • Signal processing

Background:

  • Speech enhancement (SE) algorithms can improve noisy speech intelligibility.
  • Monaural SE has historically underperformed compared to direct noisy speech training for automatic speech recognition (ASR).
  • A gap persists between SE advancements and robust ASR system development.

Purpose of the Study:

  • To bridge the gap between SE and ASR by proposing novel SE models.
  • To enable ASR systems trained solely on clean speech to perform effectively in noisy environments.
  • To advance robust ASR through improved SE frontends.

Main Methods:

  • Developed three SE models: Attentive Recurrent Network (ARN) in the time-domain, TF-CrossNet in the time-frequency domain, and MP-SENet based on magnitude-phase.
  • Decoupled SE frontend from the ASR backend, training the ASR only on clean speech.
  • Evaluated performance on WSJ, CHiME-2, LibriSpeech, and CHiME-4 corpora.

Main Results:

  • ARN, TF-CrossNet, and MP-SENet significantly improved ASR performance in noisy and reverberant conditions.
  • The proposed systems outperformed baseline ASR models trained directly on corrupted speech.
  • Achieved state-of-the-art results on CHiME-2 (5.6% WER) and CHiME-4 (3.3/4.4% WER), generalizing to real acoustic scenarios.

Conclusions:

  • The proposed SE models effectively eliminate the divide between SE and ASR.
  • Decoupled SE frontends allow for robust ASR systems trained on clean speech.
  • These advancements pave the way for more effective and generalizable robust ASR.