Related Experiment Video
Updated: Sep 11, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Improving speech emotion recognition capabilities in the short and long term using temporal bucketing and active
1Malware Lab, Cyber Security Research Center, Ben-Gurion University of the Negev, Beer-Sheva, Israel; Department of Industrial Engineering and Management, Ben-Gurion University of the Negev, Beer-Sheva, Israel.
This study introduces two novel methods for speech emotion recognition (SER), Temporal Bucketing and active learning (AL), significantly improving accuracy and reducing data labeling costs. These advancements enhance emotion inference from speech data.
Area of Science:
- Artificial Intelligence
- Data Science
- Signal Processing
Background:
- Speech Emotion Recognition (SER) has seen decades of research with various data science methods proposed.
- Improving the utilization of the temporal dimension in speech data is crucial for enhanced SER.
- Existing methods require significant data labeling efforts, impacting cost and efficiency.
Purpose of the Study:
- To introduce two novel methods for improving emotion inference from speech.
- To enhance the utilization of speech data's temporal dimension for better SER.
- To develop a cost-effective active learning (AL) method for long-term SER system improvement.
Main Methods:
- Temporal Bucketing SER: A short-term method focusing on improved utilization of speech data's temporal dimension.
- Pool-based active learning (AL) for SER: A long-term method leveraging Temporal Bucketing SER with selective sampling criteria.
- Evaluation on five diverse, publicly available datasets: RAVDESS Speech, RAVDESS Song, IEMOCAP, EMO-DB, and SAVEE.
Main Results:
- Temporal Bucketing SER outperformed state-of-the-art (SOTA) methods, achieving high accuracy rates (e.g., 91.11% on RAVDESS Speech, 95.29% on RAVDESS Song).
- The proposed AL method achieved 90% of maximum accuracy while requiring an average of 8.12% fewer samples than SOTA AL methods.
- The AL method demonstrated significant cost savings in labeling, ranging from 5% to 20% compared to SOTA AL and passive learning.
Conclusions:
- The proposed Temporal Bucketing SER method offers superior performance in emotion recognition from speech.
- The novel AL method effectively reduces the number of required labeled samples and associated costs.
- These methods provide a more efficient and accurate approach to developing and maintaining SER systems over time.
More Related Videos
Related Concept Videos
Chunking and Rehearsal in Sensory Memory
Labeling Emotion

