Related Experiment Video
Updated: Jul 2, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.5K
Exploring the performance of automatic speaker recognition using twin speech and deep learning-based artificial
Julio Cesar Cavalcanti1,2,3, Ronaldo Rodrigues da Silva4, Anders Eriksson1
1Laboratory of Phonetics, Department of Linguistics, Stockholm University, Stockholm, Sweden.
Frontiers in Artificial Intelligence
|February 26, 2024
Summary
Automatic speaker recognition (ASR) systems struggle with identical twins due to high similarity. Longer speech samples significantly improve ASR accuracy by reducing variability.
Area of Science:
- Speech processing
- Biometrics
- Machine learning
Background:
- Automatic Speaker Recognition (ASR) systems are crucial for security and authentication.
- Distinguishing between highly similar speakers, such as identical twins, remains a significant challenge for ASR systems.
- The impact of speech sample length on ASR performance, especially in challenging scenarios, requires further investigation.
Purpose of the Study:
- To evaluate the impact of speaker similarity, specifically identical twins, on ASR system performance.
- To assess the effect of speech sample length on the accuracy of speaker recognition.
- To analyze the performance of the SpeechBrain toolkit in distinguishing between highly similar speakers.
Main Methods:
- Utilized a dataset of 20 male identical twin speakers in spontaneous dialogues and interviews.
- Tested the SpeechBrain toolkit's ASR performance using speech samples ranging from 5 to 30 seconds.
- Evaluated system performance using Equal Error Rates (EER) and Log-likelihood Ratios (Cllr).
Main Results:
- Identical twins posed a substantial challenge, decreasing overall speaker recognition accuracy.
- Longer speech samples (up to 30s) resulted in significantly better performance than shorter samples.
- Increased sample size reduced variability in speaker similarity/dissimilarity scores, improving estimation.
Conclusions:
- High speaker similarity, exemplified by identical twins, significantly degrades ASR system accuracy.
- Increasing speech sample length is a viable strategy to enhance ASR performance and reduce variability.
- The degree of likeness among identical twins varies, presenting differential challenges for ASR systems.

