Validation of an embedding-based forensic voice comparison system using short speech samples in Chinese languages
Bruce Xiao Wang1, Cuiling Zhang2, Ricky K W Chan3
1Department of English and Communication, The Hong Kong Polytechnic University, Hong Kong, China.
Forensic Science International
|July 20, 2026
Summary
Forensic voice comparison performance using ECAPA-TDNN x-vectors significantly improves with speech duration up to 10 seconds. Calibration size had a lesser impact, with Hong Kong Cantonese outperforming Northeast Mandarin.
Area of Science:
- Speech processing
- Forensic science
- Machine learning for biometrics
Background:
- Speaker recognition systems are crucial in forensic voice comparison.
- ECAPA-TDNN x-vector systems offer advanced speaker embedding extraction.
- Understanding system performance under varying conditions is vital for practical application.
Purpose of the Study:
- To evaluate the performance of an ECAPA-TDNN x-vector system in forensic voice comparison.
- To investigate the impact of speech duration and calibration set size on system accuracy.
- To compare performance across different languages (Hong Kong Cantonese and Northeast Mandarin).
Main Methods:
- Utilized forensically realistic speech data from 90 Hong Kong Cantonese and 90 Northeast Mandarin speakers.
- Extracted speaker embeddings using a pre-trained ECAPA-TDNN model.
- Applied Probabilistic Linear Discriminant Analysis (PLDA) scoring and logistic regression calibration.
- Systematically varied speech duration (2-20s) and calibration group size (20-60 speakers).
Main Results:
- Performance improved significantly with speech duration up to 10 seconds, with diminishing gains thereafter.
- Hong Kong Cantonese (HKC) consistently outperformed Northeast Mandarin (NEM).
- Calibration size had a minor effect compared to duration, with more fluctuation observed with fewer than 30 speakers.
Conclusions:
- Speech duration is a critical factor for ECAPA-TDNN x-vector system performance in forensic voice comparison.
- The system demonstrates robust performance, particularly with sufficient speech duration.
- Language-specific characteristics, like content similarity, can influence comparative accuracy.
Keywords:
Chinese languagesForensic automatic speaker recognitionForensic voice comparisonLikelihood ratioMore Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
