Related Experiment Video
Updated: May 4, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Intoxicated Speech Detection: A Fusion Framework with Speaker-Normalized Hierarchical Functionals and GMM
Daniel Bone1, Ming Li1, Matthew P Black1
1Signal Analysis & Interpretation Laboratory (SAIL), University of Southern California, 3710 McClintock Ave., Los Angeles, CA 90089, USA.
Detecting speaker state from speech is challenging. This study presents a novel system using fused prosodic and spectral features, achieving high accuracy in identifying alcohol intoxication from speech signals.
Area of Science:
- Speech processing
- Computational paralinguistics
- Machine learning for behavioral analysis
Background:
- Speech signals contain rich paralinguistic information beyond linguistic content.
- Speaker state, including temporary physiological changes like alcohol intoxication, is difficult to detect computationally.
- Previous research highlights the challenge of accurately identifying speaker states from acoustic features.
Purpose of the Study:
- To develop and evaluate a robust system for detecting speaker state, specifically alcohol intoxication, from speech.
- To explore the effectiveness of fusing multiple feature representations (prosodic and spectral) for improved detection accuracy.
- To investigate methods for normalizing speaker variability to enhance system performance.
Main Methods:
- Utilized multiple representations of prosodic and spectral features.
- Employed classifier fusion techniques to combine individual model predictions.
- Investigated speaker normalization strategies, including using held-out baseline data.
- Evaluated system performance on the Alcohol Language Corpus at the Interspeech 2011 Intoxication Subchallenge.
Main Results:
- The fused system achieved an unweighted average recall (UAR) of 68.8% on the test set, outperforming the challenge baseline by 5.5% absolute.
- Speaker normalization significantly improved system performance, comparable to other normalization techniques.
- Matched-prompt training led to more consistent results, with the combined system achieving a UAR of 71.4%.
Conclusions:
- Fusing diverse feature representations and employing speaker normalization are crucial for accurate speaker state detection.
- The developed system demonstrates a significant advancement in identifying alcohol intoxication from speech.
- The findings offer a practical approach for building robust speaker state detection systems.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Extraction: Advanced Methods
Mass Spectrometry: Alcohol Fragmentation
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Expected Frequencies in Goodness-of-Fit Tests
Language and Cognition