Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

283
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
283
Air-entraining Agents01:27

Air-entraining Agents

100
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
100
Effects of feedback01:24

Effects of feedback

627
Feedback in control systems plays a critical role in shaping various operational parameters, extending beyond simple error reduction to influence stability, bandwidth, gain, impedance, and sensitivity. Understanding these effects requires examining a basic feedback system characterized by defined input, output, error, and feedback signals.
Feedback significantly modifies the gain of a control system. The gain of a system without feedback is altered by a factor of one plus GH, where G represents...
627
Double Resonance Techniques: Overview01:12

Double Resonance Techniques: Overview

252
Double resonance techniques in Nuclear Magnetic Resonance (NMR) spectroscopy involve the simultaneous application of two different frequencies or radiofrequency pulses to manipulate and observe two distinct nuclear spins. One important application of double resonance is spin decoupling, which selectively suppresses coupling with one type of nucleus while observing the NMR signal from another nucleus, simplifying the spectrum and enhancing resolution.
Spin decoupling is usually achieved by...
252
Linear Approximation in Frequency Domain01:26

Linear Approximation in Frequency Domain

120
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
120

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

On the Challenges of Acoustic Energy Mapping Using a WASN: Synchronization and Audio Capture.

Sensors (Basel, Switzerland)·2023
Same author

A Corpus-Based Evaluation of Beamforming Techniques and Phase-Based Frequency Masking.

Sensors (Basel, Switzerland)·2021
Same author

A Review on Auditory Perception for Unmanned Aerial Vehicles.

Sensors (Basel, Switzerland)·2020
Same author

On the Use of the AIRA-UAS Corpus to Evaluate Audio Processing Algorithms in Unmanned Aerial Systems.

Sensors (Basel, Switzerland)·2019

Related Experiment Video

Updated: Jul 30, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K

Characterization of Deep Learning-Based Speech-Enhancement Techniques in Online Audio Processing Applications.

Caleb Rascon1

  • 1Computer Science Department, Instituto de Investigaciones en Matematicas Aplicadas y en Sistemas, Universidad Nacional Autonoma de Mexico, Mexico City 3000, Mexico.

Sensors (Basel, Switzerland)
|May 13, 2023
PubMed
Summary

This study evaluates deep learning speech enhancement for real-time applications, analyzing performance metrics like signal-to-interference ratio, response time, and memory usage based on input length and audio conditions.

Keywords:
online applicabilityreal-time factorspeech enhancement

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

476
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

641

Related Experiment Videos

Last Updated: Jul 30, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

476
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

641

Area of Science:

  • Audio Signal Processing
  • Machine Learning
  • Digital Communications

Background:

  • Deep learning models show promise for speech enhancement in offline scenarios.
  • Evaluating these models for online, real-time audio processing is crucial for practical applications.

Purpose of the Study:

  • To assess the online applicability of state-of-the-art deep learning speech enhancement techniques.
  • To characterize model performance concerning input segment length, noise, interference, and reverberation.

Main Methods:

  • Systematic evaluation of three popular models (MetricGAN+, SFM-ML, Demucs-Denoiser) using the Speechbrain framework.
  • Analysis of output signal-to-interference ratio, response time, and memory usage.
  • Investigation of the impact of varying input lengths and audio degradation levels.

Main Results:

  • Performance metrics (e.g., signal-to-interference ratio, response time, memory usage) are influenced by input segment length.
  • Different models exhibit varying trade-offs between enhancement quality and online processing efficiency.
  • Audio conditions (noise, interference, reverberation) significantly affect model performance in online settings.

Conclusions:

  • The study provides the first comprehensive characterization of deep learning speech enhancement models for online use.
  • Recommendations for future research are proposed to optimize models for real-time digital voice communication.
  • Understanding these trade-offs is essential for selecting and deploying effective speech enhancement solutions.