Related Experiment Video
Updated: Oct 30, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Utterance Level Feature Aggregation with Deep Metric Learning for Speech Emotion Recognition
Bogdan Mocanu1, Ruxandra Tapu2, Titus Zaharia2
1Department of Telecommunications, Faculty of ETTI, University "Politehnica" of Bucharest, 060042 Bucharest, Romania.
This study presents a novel speech emotion recognition method using SE-ResNet and GhostVLAD clustering. The approach significantly improves accuracy over existing techniques and human performance.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Machine Learning
Background:
- Speech conveys paralinguistic emotional information crucial for human-computer interaction.
- Automatic Speech Emotion Recognition (SER) is vital for applications like mental health diagnosis and human behavior analysis.
- Current SER methods struggle with robust utterance-level feature representation.
Purpose of the Study:
- To introduce a novel SER method overcoming limitations of state-of-the-art techniques.
- To enhance feature representation for robust, single-utterance vector aggregation.
- To improve latent space data representation through an emotionally constrained loss function.
Main Methods:
- Utilized a Squeeze and Excitation ResNet (SE-ResNet) model with spectrogram inputs.
- Extended the CNN architecture with a trainable discriminative GhostVLAD clustering layer for feature aggregation.
- Implemented an end-to-end neural embedding approach with an emotionally constrained triplet loss function.
Main Results:
- Achieved 83.35% accuracy on the RAVDESS dataset and 64.92% on the CREMA-D dataset.
- Demonstrated accuracy gains superior to 24% compared to human observers.
- Showcased accuracy improvements of over 3% against state-of-the-art methods.
Conclusions:
- The proposed SE-ResNet and GhostVLAD-based SER method offers robust utterance-level feature representation.
- The emotionally constrained triplet loss function effectively improves latent space representation.
- The methodology significantly advances the state-of-the-art in automatic speech emotion recognition.
More Related Videos
06:37Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
Related Concept Videos
Labeling Emotion
Sound Intensity Level
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Non-Verbal Cues