Related Experiment Video
Updated: Nov 10, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
On the Speech Properties and Feature Extraction Methods in Speech Emotion Recognition
Juraj Kacur1, Boris Puterka2, Jarmila Pavlovicova2
1Institute of Multimedia Information and Communication Technologies, Faculty of Electrical Engineering and Information Technology, Slovak University of Technology in Bratislava, 2412 Bratislava, Slovakia.
This study reveals how speech characteristics and processing methods impact speech emotion recognition accuracy. Vocal tract features using psychoacoustic filter banks achieved 75% accuracy for seven emotions.
Area of Science:
- Speech processing and recognition
- Computational linguistics
- Affective computing
Background:
- Existing speech emotion recognition (SER) systems lack detailed understanding of how specific acoustic features and processing choices influence performance.
- A need exists to analyze fundamental speech characteristics and modeling techniques for improved SER system design.
Purpose of the Study:
- To extend the physical perspective on SER by analyzing basic speech characteristics and modeling methods.
- To investigate the impact of various signal processing techniques on SER accuracy.
- To identify optimal settings for robust emotion recognition from speech.
Main Methods:
- Analysis of time characteristics (segmentation, windowing), frequency ranges/scales, spectrograms, vocal tract modeling (filter banks, LPC), and excitation signals.
- Evaluation of cepstral features, magnitude/phase manipulations, and advanced classification methods.
- Application of rigorous statistical tests including N-fold cross-validation, paired t-tests, rank, and Pearson correlations.
Main Results:
- Several settings achieved approximately 75% accuracy in recognizing seven distinct emotions.
- Vocal tract features utilizing psychoacoustic filter banks (0-8 kHz) yielded the most successful results.
- Spectrograms containing vocal tract and excitation information also performed well.
- Basic processing steps like pre-emphasis and segmentation significantly influenced outcomes.
Conclusions:
- The study provides crucial insights into the physical underpinnings of SER, guiding future system development.
- Optimal performance is linked to specific vocal tract feature extraction and processing parameters.
- Even elementary signal processing choices have a substantial effect on SER accuracy, with findings demonstrating robustness across datasets.
More Related Videos
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Labeling Emotion
Non-Verbal Cues
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Physiology of Emotion
Autonomic Nervous System
The autonomic nervous system (ANS) plays a critical role in emotional responses by regulating involuntary physiological functions. It consists of two main components: the sympathetic and parasympathetic systems. The sympathetic system...
Cognitive Theories: Schachter-Singer Theory of Emotion
Physiological Arousal and Cognitive Labeling
According to this theory, when an individual experiences...
Role of Emotions in Social Life