Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

205
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
205
Chunking and Rehearsal in Sensory Memory01:22

Chunking and Rehearsal in Sensory Memory

198
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
198
Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

105
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
105

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Multimodal Environmental Sensing Using AI & IoT Solutions: A Cognitive Sound Analysis Perspective.

Sensors (Basel, Switzerland)·2024
Same author

Development and evaluation of a tablet-based diagnostic audiometer.

International journal of audiology·2019
See all related articles

Related Experiment Video

Updated: Jun 21, 2025

Testing Sensory and Multisensory Function in Children with Autism Spectrum Disorder
09:13

Testing Sensory and Multisensory Function in Children with Autism Spectrum Disorder

Published on: April 22, 2015

16.5K

Multisensory Fusion for Unsupervised Spatiotemporal Speaker Diarization.

Paris Xylogiannis1, Nikolaos Vryzas1, Lazaros Vrysis1

  • 1Multidisciplinary Media & Mediated Communication Research Group (M3C), Aristotle University, 54636 Thessaloniki, Greece.

Sensors (Basel, Switzerland)
|July 13, 2024
PubMed
Summary

This study enhances speaker diarization by integrating speaker embeddings with spatial Time Difference of Arrival (TDOA) data. Combining these features significantly improves accuracy in identifying "who spoke when" in meeting recordings.

Keywords:
AI-enabled systemsdeep learningmultimodal decision makingsmartphonessound localizationspeaker diarization

More Related Videos

Using Informational Connectivity to Measure the Synchronous Emergence of fMRI Multi-voxel Information Across Time
07:12

Using Informational Connectivity to Measure the Synchronous Emergence of fMRI Multi-voxel Information Across Time

Published on: July 1, 2014

12.3K
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K

Related Experiment Videos

Last Updated: Jun 21, 2025

Testing Sensory and Multisensory Function in Children with Autism Spectrum Disorder
09:13

Testing Sensory and Multisensory Function in Children with Autism Spectrum Disorder

Published on: April 22, 2015

16.5K
Using Informational Connectivity to Measure the Synchronous Emergence of fMRI Multi-voxel Information Across Time
07:12

Using Informational Connectivity to Measure the Synchronous Emergence of fMRI Multi-voxel Information Across Time

Published on: July 1, 2014

12.3K
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K

Area of Science:

  • Speech processing
  • Machine learning
  • Acoustic signal processing

Background:

  • Speaker diarization identifies speakers in audio recordings.
  • Spatial features from microphone arrays can aid speaker identification in meetings.
  • Current methods often rely solely on audio features, potentially missing spatial cues.

Purpose of the Study:

  • To develop and evaluate a framework combining speaker embeddings with Time Difference of Arrival (TDOA) for improved speaker diarization.
  • To assess the effectiveness of different speaker embedding models (ECAPA-TDNN, X-vectors) when integrated with spatial information.
  • To determine optimal methods for combining spatial-temporal information for clustering speakers.

Main Methods:

  • Speaker embeddings extracted using ECAPA-TDNN and X-vectors.
  • Time Difference of Arrival (TDOA) calculated using Generalized Cross-Correlation (GCC) with Phase Transform (PHAT) weighting.
  • Framework evaluated on AVLab and SpeaD-M3C multichannel datasets, assessing various spatial-temporal combination techniques.

Main Results:

  • Integration of spatial information significantly improves state-of-the-art deep learning diarization models.
  • A 2-3% reduction in Diarization Error Rate (DER) was observed compared to baseline approaches.
  • ECAPA-TDNN generally outperformed X-vectors, but both benefited from spatial data integration.

Conclusions:

  • Combining speaker embeddings with spatial TDOA data is an effective strategy for enhancing speaker diarization performance.
  • Spatial information provides valuable cues that complement audio-based speaker recognition, especially in meeting environments.
  • The proposed framework demonstrates the utility of leveraging multichannel audio for more accurate and robust speaker identification.