Related Experiment Video
Updated: Aug 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Cross-Corpus Speech Emotion Recognition Based on Transfer Learning and Multi-Loss Dynamic Adjustment.
Huawei Tao1, Yang Wang1, Zhihao Zhuang1
1College of Information Science and Engineering, Henan University of Technology, Zhengzhou 450001, China.
This study introduces a novel algorithm for cross-corpus speech emotion recognition (SER). The transfer learning and multi-loss dynamic adjustment (TLMLDA) method enhances feature representation and generalization for improved accuracy.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Speech Processing
Background:
- Cross-corpus speech emotion recognition (SER) faces challenges due to mismatched feature distributions between training and testing datasets.
- Classical algorithms often struggle to achieve optimal performance in cross-corpus SER scenarios.
Purpose of the Study:
- To propose an effective algorithm for cross-corpus SER that addresses feature distribution mismatches.
- To improve the generalization ability and recognition accuracy of SER systems when applied to unseen speech corpora.
Main Methods:
- A novel deep network model combining a deep auto-encoder and fully connected layers was developed to enhance feature representation.
- Global and subdomain adaptive algorithms were employed for effective feature transfer across different speech corpora.
- Dynamic weighting factors were introduced to balance multiple loss functions, preventing optimization offset during model training.
Main Results:
- The proposed transfer learning and multi-loss dynamic adjustment (TLMLDA) algorithm demonstrated excellent recognition results on the Berlin, eNTERFACE, and CASIA speech corpora.
- The algorithm achieved competitive performance compared to most state-of-the-art methods in cross-corpus SER tasks.
- The TLMLDA algorithm effectively improved the generalization ability of the SER system.
Conclusions:
- The TLMLDA algorithm offers a robust solution for cross-corpus speech emotion recognition by effectively handling feature distribution discrepancies.
- The proposed method shows significant potential for real-world applications requiring accurate emotion detection from diverse speech data.
- This research contributes to advancing the field of SER by providing a more generalizable and accurate approach.
More Related Videos
Related Concept Videos
Labeling Emotion
Multi-input and Multi-variable systems
In the absence...
Improving Translational Accuracy
Physiology of Emotion
Autonomic Nervous System
The autonomic nervous system (ANS) plays a critical role in emotional responses by regulating involuntary physiological functions. It consists of two main components: the sympathetic and parasympathetic systems. The sympathetic system...
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Facial Feedback Hypothesis

