Related Experiment Video
Updated: Sep 17, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
MS-EmoBoost: a novel strategy for enhancing self-supervised speech emotion representations.
Hongchen Song1, Long Zhang2, Meixian Gao1
1College of Computer and Information Engineering, Tianjin Normal University, Tianjin, 300387, China.
This study introduces MS-EmoBoost to improve speech emotion recognition (SER) by enhancing self-supervised learning (SSL) features. The method boosts emotional representation, achieving competitive accuracy on benchmark datasets.
Area of Science:
- Speech Emotion Recognition (SER)
- Machine Learning
- Signal Processing
Background:
- Accurate Speech Emotion Recognition (SER) relies on rich emotional representations from raw speech.
- Self-supervised learning (SSL) shows promise for SER feature extraction, mirroring success in Automatic Speech Recognition (ASR).
- Current SSL methods lack sufficient sensitivity to emotional nuances, limiting their effectiveness in SER.
Purpose of the Study:
- To propose MS-EmoBoost, a novel strategy for enhancing self-supervised speech emotion representations.
- To improve the emotional representation capabilities of SSL features for SER tasks.
- To validate the effectiveness of MS-EmoBoost on benchmark SER datasets.
Main Methods:
- MS-EmoBoost leverages deep emotional information from Melfrequency cepstral coefficients (MFCC) and spectrograms.
- This emotional guidance enhances the capabilities of self-supervised features for emotion recognition.
- The approach was tested using wav2vec 2.0 Base features on IEMOCAP, EMODB, and EMOVO datasets.
Main Results:
- MS-EmoBoost significantly enhanced the emotional representation of wav2vec 2.0 Base features.
- The method achieved competitive SER performance across three benchmark datasets: IEMOCAP (WA:72.10%, UA:72.91%), EMODB (WA:92.45%, UA:92.62%), and EMOVO (WA:86.88%, UA:87.51%).
- The strategy proved effective for other self-supervised features as well.
Conclusions:
- MS-EmoBoost offers a viable strategy for improving SER by enhancing SSL features.
- The proposed method effectively captures and utilizes emotional information for more accurate speech emotion recognition.
- This approach advances the field of SER by making SSL methods more sensitive to emotional content.
More Related Videos
Related Concept Videos
Labeling Emotion
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Coping Strategies: Emotion Focused
Self-Evaluation: Self-Enhancement and Self-Verification
Physiology of Emotion
Autonomic Nervous System
The autonomic nervous system (ANS) plays a critical role in emotional responses by regulating involuntary physiological functions. It consists of two main components: the sympathetic and parasympathetic systems. The sympathetic system...
Motional Emf

