Related Experiment Video
Updated: Feb 26, 2026

09:09
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
929
Detecting paralinguistic events in audio stream using context in features and probabilistic decisions.
Rahul Gupta1, Kartik Audhkhasi2, Sungbok Lee1
1Signal Analysis and Interpretation Laboratory (SAIL), University of Southern California, 3710 McClintock Avenue, Los Angeles, CA 90089, USA.
Computer Speech & Language
|July 18, 2017
Summary
This study presents an automatic system for detecting non-verbal vocal events like laughter and fillers in speech. The system achieves high accuracy, improving conversational analysis through enhanced non-verbal cue detection.
Area of Science:
- Speech processing and human-computer interaction
- Computational linguistics and affective computing
- Machine learning for signal processing
Background:
- Non-verbal communication, including vocal cues like laughter and fillers, is crucial for conversational flow, emotion expression, and personality marking.
- Automatic detection of these non-verbal vocal events is challenging but valuable for understanding human interaction.
- Previous work established a baseline system in the Interspeech 2013 Social Signals Sub-challenge.
Purpose of the Study:
- To develop and enhance an automatic system for detecting non-verbal vocal events, specifically laughter and fillers.
- To investigate the incorporation of local context at feature and decision levels for improved detection accuracy.
- To analyze feature sensitivity and its impact on system design for non-verbal vocal event detection.
Main Methods:
- Extension of a previous winning system from Interspeech 2013 for frame-wise event detection.
- Incorporation of local context at two levels: raw frame-wise features and output decisions.
- Application of heuristic rules to refine frame-based predictions and reduce errors.
Main Results:
- Achieved an Area Under the ROC curve of 95.3% for laughter detection.
- Achieved an Area Under the ROC curve of 90.4% for filler detection on the Interspeech 2013 test set.
- Feature sensitivity analysis indicated that features with higher discriminability positively impact system performance.
Conclusions:
- The developed system effectively detects laughter and fillers with high accuracy, outperforming previous benchmarks.
- Incorporating local context and heuristic post-processing significantly enhances the robustness of non-verbal vocal event detection.
- Understanding feature importance is key for designing more effective and sensitive automatic speech analysis systems.
Related Concept Videos
Non-Verbal Cues
384
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
384
Perceiving Loudness, Pitch, and Location
1.2K
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
1.2K

