Related Experiment Video
Updated: Jul 19, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
End-to-end neural speaker diarization with an iterative adaptive attractor estimation.
Fengyuan Hao1, Xiaodong Li1, Chengshi Zheng1
1Key Laboratory of Noise and Vibration Research, Institute of Acoustics, Chinese Academy of Sciences, Beijing, 100190, China; University of Chinese Academy of Sciences, Beijing, 100049, China.
This study introduces an iterative adaptive attractor estimation network to improve end-to-end neural diarization (EEND) performance. The novel approach refines speaker diarization results, significantly reducing errors on simulated and real-world datasets.
Area of Science:
- Speech processing
- Machine learning
- Artificial intelligence
Background:
- End-to-end neural diarization (EEND) shows promise for speaker diarization and overlapping speech, but fixed attractors limit generalization.
- Existing EEND methods struggle with unseen data due to limitations in estimating speaker-specific speech activities.
Purpose of the Study:
- To enhance the generalization and accuracy of EEND speaker diarization.
- To introduce an iterative adaptive attractor estimation (IAAE) network for refining diarization results.
Main Methods:
- Developed an IAAE network utilizing self-attentive EEND (SA-EEND) for initialization.
- Incorporated an attention-based pooling mechanism for iterative attractor estimation.
- Employed transformer decoder blocks for adaptive attractor calculation and a unified training framework.
Main Results:
- Achieved relative reductions in Diarization Error Rate (DER) of up to 44.8% on simulated data and 23.6% on the CALLHOME dataset at the second iteration.
- Demonstrated further DER reduction to 7.36% on the CALLHOME dataset with increased refinement steps.
- Outperformed the baseline SA-EEND method, showcasing improved diarization accuracy.
Conclusions:
- The proposed IAAE network significantly improves EEND speaker diarization performance and generalization.
- Iterative refinement and adaptive attractors lead to more discriminable speaker embeddings and state-of-the-art results.
- The method offers a robust solution for accurate speaker diarization in complex acoustic environments.

