Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Non-Verbal Cues01:29

Non-Verbal Cues

Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A speech prediction model based on codec modeling and transformer decoding.

Computer speech & languageยท2026
Same author

A Molecular Trimming Strategy for Hypoxia-Tolerant Photosensitizers With Enhanced cGAS-STING Activation.

Angewandte Chemie (International ed. in English)ยท2026
Same author

Towards decoupling frontend enhancement and backend recognition in monaural robust ASR.

Computer speech & languageยท2026
Same author

Efficacy of SWIM technology combined with direct aspiration first pass technique for large vessel occlusion in acute ischemic stroke.

American journal of translational researchยท2026
Same author

Manipulating RTP properties of the same organic molecule by polymorphic engineering.

Chemical communications (Cambridge, England)ยท2025
Same author

Confined Growth of 2D Covalent Organic Framework Nanosheets with Controlled Thickness for Osmotic Energy Conversion.

Small (Weinheim an der Bergstrasse, Germany)ยท2025

Related Experiment Video

Updated: Jul 7, 2026

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
07:52

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners

Published on: March 13, 2026

Monaural speech segregation based on pitch tracking and amplitude modulation.

Guoning Hu1, Deliang Wang

  • 1Biophys. Program, Ohio State Univ., Columbus, OH, USA.

IEEE Transactions on Neural Networks
|February 2, 2008
PubMed
Summary

This study introduces a new method for separating voiced speech from a single audio source. The novel system effectively segregates high-frequency speech components by treating resolved and unresolved harmonics differently.

More Related Videos

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Related Experiment Videos

Last Updated: Jul 7, 2026

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
07:52

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners

Published on: March 13, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
05:48

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis

Published on: August 9, 2024

Area of Science:

  • Speech processing
  • Auditory perception
  • Signal processing

Background:

  • Monaural speech segregation is challenging, particularly for high-frequency speech components.
  • Existing systems struggle with high-frequency speech due to limitations in auditory scene analysis.
  • Psychoacoustic evidence indicates distinct mechanisms for processing resolved and unresolved harmonics.

Purpose of the Study:

  • To propose a novel system for voiced speech segregation that addresses limitations in current methods.
  • To develop a system that segregates resolved and unresolved harmonics using different strategies.
  • To improve the segregation of high-frequency speech components.

Main Methods:

  • The proposed system segregates resolved harmonics based on temporal continuity and cross-channel correlation.
  • Unresolved harmonics are segregated using common amplitude modulation (AM) and temporal continuity.
  • A pitch contour, estimated from dominant pitch and refined by psychoacoustic constraints, guides the segregation process.

Main Results:

  • The novel system demonstrates substantially improved performance compared to previous methods.
  • Performance gains are particularly significant for the high-frequency components of speech.
  • The differential treatment of resolved and unresolved harmonics leads to better overall segregation.

Conclusions:

  • The proposed system offers a significant advancement in monaural voiced speech segregation.
  • Differentiating the processing of resolved and unresolved harmonics is crucial for effective speech separation.
  • The system shows promise for applications requiring robust speech extraction from complex acoustic environments.