Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Interaction Between Head Movement Behavior and Simulated Spatial Filtering: Comparing Free Conversation and Speech Tests in Virtual Reality.

Trends in hearing·2026
Same author

Effect of Avatar Head Movements on Communication Behavior and Subjective Evaluations of Presence and Success in Triadic Conversations.

Trends in hearing·2026
Same author

Neural speech tracking in a virtual acoustic environment: audio-visual benefit for unscripted continuous speech.

Frontiers in human neuroscience·2025
Same author

Benefit of Hearing-Aid Amplification and Signal Enhancement for Speech Reception in Complex Listening Situations.

Trends in hearing·2024
Same author

The future of hearing aid technology : Can technology turn us into superheroes?

Zeitschrift fur Gerontologie und Geriatrie·2023
Same author

Vehicle noise: comparison of loudness ratings in the field and the laboratory.

International journal of audiology·2022

Related Experiment Video

Updated: Feb 25, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.0K

Modeling speech localization, talker identification, and word recognition in a multi-talker setting.

Angela Josupeit1, Volker Hohmann1

  • 1Medizinische Physik and Cluster of Excellence Hearing4all, Universität Oldenburg, 26111 Oldenburg, Germany.

The Journal of the Acoustical Society of America
|August 3, 2017
PubMed
Summary

This study presents an auditory model that decodes complex multi-talker scenes using extracted speech features called "glimpses." The model accurately predicts human performance in tasks like target identification and word recognition.

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

929
Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention
04:32

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention

Published on: December 20, 2024

935

Related Experiment Videos

Last Updated: Feb 25, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.0K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

929
Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention
04:32

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention

Published on: December 20, 2024

935

Area of Science:

  • Auditory Neuroscience
  • Computational Auditory Scene Analysis
  • Psychoacoustics

Background:

  • Understanding auditory scene analysis in multi-talker environments is challenging.
  • Previous models often require complex source superposition assumptions.

Purpose of the Study:

  • Introduce a novel computational model for auditory tasks in multi-talker settings.
  • Simulate and validate the model against psychoacoustic data from human listeners.
  • Investigate the sufficiency of sparse auditory features for scene decoding.

Main Methods:

  • Developed a model extracting salient auditory features ('glimpses') from multi-talker signals.
  • Utilized periodicity, periodic energy, and interaural time/level differences as key features.
  • Employed a classification method comparing target templates to extracted glimpses.

Main Results:

  • Model performance significantly exceeded chance levels across all auditory tasks (localization, identification, recognition).
  • Model predictions showed strong agreement with human subject data.
  • Demonstrated that sparse 'glimpses' contain adequate information for complex auditory scene analysis.

Conclusions:

  • Sparse auditory features ('glimpses') are sufficient for decoding complex multi-talker scenes.
  • Complex source superposition models may not be necessary for auditory scene analysis.
  • Simple models of clean speech features can decode challenging multi-talker environments.