Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Elaborative Rehearsals01:07

Elaborative Rehearsals

181
Elaborative rehearsal is a crucial cognitive strategy that strengthens information encoding in long-term memory by making meaningful connections between new data and pre-existing knowledge. This approach contrasts with maintenance rehearsal, which involves simple repetition without delving into the significance of the information. While maintenance rehearsal might temporarily keep information active in short-term memory, it is less effective for long-term retention.
The effectiveness of...
181
Double Resonance Techniques: Overview01:12

Double Resonance Techniques: Overview

478
Double resonance techniques in Nuclear Magnetic Resonance (NMR) spectroscopy involve the simultaneous application of two different frequencies or radiofrequency pulses to manipulate and observe two distinct nuclear spins. One important application of double resonance is spin decoupling, which selectively suppresses coupling with one type of nucleus while observing the NMR signal from another nucleus, simplifying the spectrum and enhancing resolution.
Spin decoupling is usually achieved by...
478
Echo01:06

Echo

696
The human ear cannot distinguish between two sources of sound if they happen to reach within a specific time interval, typically 0.1 seconds apart. More than this, and they are perceived as separate sources.
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
696
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

625
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
625
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

8.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.3K
Chunking and Rehearsal in Sensory Memory01:22

Chunking and Rehearsal in Sensory Memory

390
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
390

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A speech prediction model based on codec modeling and transformer decoding.

Computer speech & language·2026
Same author

A Molecular Trimming Strategy for Hypoxia-Tolerant Photosensitizers With Enhanced cGAS-STING Activation.

Angewandte Chemie (International ed. in English)·2026
Same author

Towards decoupling frontend enhancement and backend recognition in monaural robust ASR.

Computer speech & language·2026
Same author

Efficacy of SWIM technology combined with direct aspiration first pass technique for large vessel occlusion in acute ischemic stroke.

American journal of translational research·2026
Same author

Manipulating RTP properties of the same organic molecule by polymorphic engineering.

Chemical communications (Cambridge, England)·2025
Same author

Confined Growth of 2D Covalent Organic Framework Nanosheets with Controlled Thickness for Osmotic Energy Conversion.

Small (Weinheim an der Bergstrasse, Germany)·2025

Related Experiment Video

Updated: Nov 12, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K

Deep Learning for Talker-dependent Reverberant Speaker Separation: An Empirical Study.

Masood Delfarah1, DeLiang Wang1

  • 1Computer Science and Engineering, The Ohio State University, Columbus, OH, USA.

IEEE/ACM Transactions on Audio, Speech, and Language Processing
|March 22, 2021
PubMed
Summary

This study tackles speaker separation in real-world reverberant conditions using bidirectional long short-term memory (BLSTM) networks. A novel two-stage approach significantly improves speech separation and dereverberation compared to existing methods.

Keywords:
Cochannel speech separationdeep neural networksspeech dereverberationtwo-stage network

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

653
Investigating the Three-dimensional Flow Separation Induced by a Model Vocal Fold Polyp
09:58

Investigating the Three-dimensional Flow Separation Induced by a Model Vocal Fold Polyp

Published on: February 3, 2014

8.7K

Related Experiment Videos

Last Updated: Nov 12, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

653
Investigating the Three-dimensional Flow Separation Induced by a Model Vocal Fold Polyp
09:58

Investigating the Three-dimensional Flow Separation Induced by a Model Vocal Fold Polyp

Published on: February 3, 2014

8.7K

Area of Science:

  • Signal Processing
  • Artificial Intelligence

Background:

  • Speaker separation is crucial for isolating speech from mixed audio.
  • Existing methods often fail in reverberant environments, limiting real-world applications.

Purpose of the Study:

  • To develop a talker-dependent speaker separation system for reverberant conditions.
  • To improve speech clarity and reduce reverberation in mixed audio signals.

Main Methods:

  • Utilized recurrent neural networks with bidirectional long short-term memory (BLSTM).
  • Proposed a two-stage network architecture for separation and dereverberation.
  • Employed time-frequency masking for enhanced performance.

Main Results:

  • The two-stage BLSTM network significantly improved speaker separation and dereverberation.
  • Achieved substantial gains over unprocessed mixtures and single-stage networks.
  • Demonstrated superior performance of time-frequency masking over spectral mapping.

Conclusions:

  • The proposed two-stage BLSTM network is effective for speaker separation in reverberant environments.
  • This approach offers a significant advancement for real-world speech processing applications.
  • Time-frequency masking is a key technique for improving reverberant speaker separation.