Validation of an embedding-based forensic voice comparison system using short speech samples in Chinese languages
Bruce Xiao Wang1, Cuiling Zhang2, Ricky K W Chan3
1Department of English and Communication, The Hong Kong Polytechnic University, Hong Kong, China.
Abstract:
The current study evaluates the performance of an ECAPA-TDNN x-vector system in forensic voice comparison as a function of speech duration and calibration size. Forensically realistic speech from 90 Hong Kong Cantonese (HKC) and 90 Northeast Mandarin (NEM) speakers were used. Speaker embeddings were extracted from a pre-trained ECAPA-TDNN embedding model, followed by PLDA scoring and logistic regression calibration. System performance was evaluated with speech duration increasing from 2 to 20 s with a one-second increase and calibration group size from 20 to 60 speakers with a 10-speaker increase. Across both datasets, performance improved markedly as duration increased up to 10 s; gains beyond 10 s were modest. For HKC, the best result was achieved with 19 s and 40 calibration speakers (Cllr = 0.112), while the worst was at 2 s and 20 speakers (Cllr = 0.702). For NEM, the best was 20 s and 60 speakers (Cllr = 0.172), with the worst at 2 s and 20 speakers (Cllr = 0.884). The number of calibration speakers had a minor effect relative to duration; fluctuation in Cllr cal was more evident with fewer than 30 speakers, especially for short samples. HKC consistently outperformed NEM, likely due to greater content similarity across participants. Results were discussed in relation to previous studies using other languages.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
