Related Experiment Video
Updated: Jul 15, 2026

12:43
A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
34.7K
Fusing linguistic and acoustic information for automated forensic speaker comparison.
E K Sergidou1, Rolf Ypma2, Johan Rohdin3
1Netherlands Forensic Institute, PO Box 24044, 2490 AA The Hague, the Netherlands; University of Amsterdam, Science Park 904, 1098 XH Amsterdam, the Netherlands.
Science & Justice : Journal of the Forensic Science Society
|September 14, 2024
Summary
Combining acoustic and linguistic analysis improves speaker verification accuracy, especially for noisy phone calls. This fusion approach enhances forensic evidence reporting within the likelihood ratio framework.
Area of Science:
- Forensic Science
- Computational Linguistics
- Biometrics
Background:
- Speaker verification is critical for forensic attribution of speech.
- Current methods rely on auditory, acoustic, and automated systems using deep neural networks.
- Linguistic analysis, particularly frequent word usage, offers complementary information.
Purpose of the Study:
- To investigate the fusion of acoustic and frequent word-based linguistic analysis for speaker verification.
- To determine the optimal methods for combining these analyses within the likelihood ratio framework.
- To evaluate the effectiveness of fusion under various conditions, including noisy data.
Main Methods:
- Developed three fusion approaches: Support Vector Machine (SVM), bivariate normal distributions, and acoustic score integration.
- Utilized the FRIDA dataset and FISHER corpus for method application.
- Evaluated performance using log likelihood ratio cost (C_llr) and equal error rate (EER).
Main Results:
- Fusion of acoustic and linguistic features can significantly improve speaker verification performance.
- The benefits of fusion are particularly pronounced in challenging conditions, such as intercepted phone calls with background noise.
- The choice of fusion method impacts the overall effectiveness.
Conclusions:
- Combining modern acoustic systems with frequent word analysis offers a valuable enhancement for speaker verification.
- The likelihood ratio framework effectively integrates evidence from both acoustic and linguistic modalities.
- This integrated approach strengthens the reliability of forensic speaker identification, especially in real-world scenarios.
Keywords:
Forensic speaker comparisonFrequent-word analysisInformation fusionLikelihood ratio frameworkMulti-modal analysis
