Related Experiment Videos
Effects of applying Gaussianising transformations to DNN embeddings before calculating likelihood ratios
Rafael Oliveira Ribeiro1, Philip Weber2, Geoffrey Stewart Morrison3
1Forensic Data Science Laboratory, Aston University, Birmingham, UK; National Institute of Criminalistics, Brazilian Federal Police, Brasília, Brazil.
Abstract:
This paper investigates the effects of applying multiple different sequences of transformations that aim to Gaussianise the distribution of deep-neural-network embeddings (DNN embeddings) before those DNN embeddings are used to calculate likelihood ratios using a model that assumes Gaussian distributions (Probabilistic Linear Discriminant Analysis, PLDA). The transformations were originally developed for non-forensic applications of automatic speaker recognition. This paper investigates the effects of applying the transformations under three different sets of forensically realistic conditions. The paper assesses the effects of the transformations with respect to theoretical properties of high-dimensional Gaussian distributions, and with respect to their impact on the performance of a forensic-voice-comparison system. Considering both the distribution and the performance criteria, the sequence of transformations that appeared to give the best results was: linear discriminant functions + centring + whitening + radial Gaussianisation. The results are potentially generalisable and not just specific to forensic voice comparison, i.e., they are also potentially applicable to use of DNN embedding to calculate likelihood ratios in other branches of forensic science such as forensic comparison of facial images.