Related Experiment Video
Updated: Sep 19, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Feature and classifier-level domain adaptation in DistilHuBERT for cross-corpus speech emotion recognition.
Niloufar Naeeni1, Babak Nasersharif1
1Computer Engineering Department, K. N. Toosi University of Technology, Shariati Ave., Tehran, Iran.
This study enhances cross-corpus speech emotion recognition using DistilHuBERT and domain adaptation. The best method achieved 92.01% accuracy, improving emotion identification across diverse datasets.
Area of Science:
- Speech processing and machine learning
- Artificial intelligence in affective computing
Background:
- Cross-corpus speech emotion recognition (CCSER) faces challenges due to dataset variations like gender and language.
- Robust models are needed to accurately identify speaker emotions across diverse datasets.
Purpose of the Study:
- To develop effective domain adaptation strategies for CCSER using self-supervised speech representations.
- To improve the accuracy and robustness of emotion recognition models across different speech datasets.
Main Methods:
- Utilized DistilHuBERT for self-supervised speech representations.
- Proposed four Feature-level Domain Adaptation (FDA) methods, including Siamese networks and fine-tuning DistilHuBERT layers.
- Implemented Classifier-level Domain Adaptation (CDA) by updating the classifier with target dataset portions.
Main Results:
- The fourth FDA method, combined with CDA, yielded the highest accuracy.
- Achieved 92.01% accuracy on the EMODB dataset using ShEMO as the source dataset.
- Demonstrated the effectiveness of proposed domain adaptation techniques in CCSER.
Conclusions:
- The proposed FDA and CDA methods significantly improve CCSER performance.
- Fine-tuning both CNN and transformer layers of DistilHuBERT (fourth FDA method) combined with CDA is a promising approach for robust emotion recognition.
- This research contributes to more accurate and reliable emotion identification systems across varied speech data.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
Related Concept Videos
Labeling Emotion
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Systems-II
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as: