Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Neural Regulation01:37

Neural Regulation

41.1K
Digestion begins with a cephalic phase that prepares the digestive system to receive food. When our brain processes visual or olfactory information about food, it triggers impulses in the cranial nerves innervating the salivary glands and stomach to prepare for food.
41.1K
Downsampling01:20

Downsampling

355
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
355
Neural Circuits01:25

Neural Circuits

2.1K
Neural circuits and neuronal pools are two of the main structures found in the nervous system. Neural circuits are networks of neurons that work together to carry out a specific task or process. They consist of interconnected neurons and glial cells, which provide structural and metabolic support.
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
2.1K
Air-entraining Agents01:27

Air-entraining Agents

159
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
159
Linear Approximation in Frequency Domain01:26

Linear Approximation in Frequency Domain

226
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
226
Buffer Effectiveness02:19

Buffer Effectiveness

52.3K
Buffer solutions do not have an unlimited capacity to keep the pH relatively constant . Instead, the ability of a buffer solution to resist changes in pH relies on the presence of appreciable amounts of its conjugate weak acid-base pair. When enough strong acid or base is added to substantially lower the concentration of either member of the buffer pair, the buffering action within the solution is compromised.
The buffer capacity is the amount of acid or base that can be added to a given volume...
52.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Rare Pituitary Tumor.

Cureus·2024
Same author

Robust Three-Microphone Speech Source Localization Using Randomized Singular Value Decomposition.

IEEE access : practical innovations, open solutions·2021
Same author

Smartphone-based single-channel speech enhancement application for hearing aids.

The Journal of the Acoustical Society of America·2021
Same author

Spectral Flux-Based Convolutional Neural Network Architecture for Speech Source Localization and Its Real-Time Implementation.

IEEE access : practical innovations, open solutions·2021
Same author

CONVOLUTIONAL RECURRENT NEURAL NETWORK BASED DIRECTION OF ARRIVAL ESTIMATION METHOD USING TWO MICROPHONES FOR HEARING STUDIES.

IEEE International Workshop on Machine Learning for Signal Processing : [proceedings]. IEEE International Workshop on Machine Learning for Signal Processing·2021
Same author

Behavioral Validation of the Smartphone for Remote Microphone Technology.

Seminars in hearing·2020

Related Experiment Video

Updated: Nov 8, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K

Real-time single-channel deep neural network-based speech enhancement on edge devices.

Nikhil Shankar1, Gautam Shreedhar Bhat1, Issa M S Panahi1

  • 1Department of Electrical and Computer Engineering, The University of Texas at Dallas, Richardson, TX-75080, USA.

Interspeech
|April 26, 2021
PubMed
Summary

This study introduces a novel deep neural network for real-time speech enhancement (SE) on smartphones. The model significantly improves noisy speech quality and intelligibility, outperforming existing methods.

Keywords:
neural networksreal-timesmartphonespeech enhancement

More Related Videos

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

749
Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
06:24

Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications

Published on: January 5, 2024

1.1K

Related Experiment Videos

Last Updated: Nov 8, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

749
Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
06:24

Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications

Published on: January 5, 2024

1.1K

Area of Science:

  • Artificial Intelligence
  • Signal Processing
  • Acoustics

Background:

  • Speech enhancement (SE) is crucial for clear audio communication.
  • Existing methods often struggle with real-time processing on edge devices.
  • Deep learning offers potential for advanced SE algorithms.

Purpose of the Study:

  • To develop a real-time, single-channel speech enhancement model using deep neural networks.
  • To implement and evaluate the model on a smartphone for practical usability.
  • To compare the proposed method against conventional and deep learning-based SE techniques.

Main Methods:

  • A hybrid deep neural network architecture combining convolutional neural network (CNN) and recurrent neural network (RNN) layers was designed.
  • The model processes the noisy speech magnitude spectrum on a frame-by-frame basis.
  • Implementation on a smartphone (edge device) for real-time performance evaluation.

Main Results:

  • The proposed deep learning model demonstrated effective real-time speech enhancement on a smartphone.
  • Objective metrics, including Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI), were used for comparison.
  • Subjective listening tests confirmed superior performance compared to baseline SE methods.

Conclusions:

  • The developed CNN-RNN architecture provides a viable solution for real-time speech enhancement on edge devices.
  • The model achieves significant improvements in speech quality and intelligibility.
  • This research paves the way for enhanced mobile audio experiences.