Related Experiment Video
Updated: Sep 19, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Feature and classifier-level domain adaptation in DistilHuBERT for cross-corpus speech emotion recognition
Niloufar Naeeni1, Babak Nasersharif1
1Computer Engineering Department, K. N. Toosi University of Technology, Shariati Ave., Tehran, Iran.
None:
Cross-corpus speech emotion recognition (CCSER) aims to develop robust models capable of accurately identifying a speaker's emotional state across diverse datasets. This task is challenged by variations in dataset characteristics, such as differences in gender distribution and languages. Overcoming these challenges necessitates novel feature representation and classification strategies. This study utilizes self-supervised speech representations generated by the DistilHuBERT model and proposes domain adaptation methods at both the Feature-level (FDA) and Classifier-level (CDA) for CCSER. This paper proposes four FDA methods for adapting Cross-Corpus Speech Emotion Recognition to the target dataset. The first method integrates the CNN encoder of DistilHuBERT into a Siamese network with a contrastive loss function to learn a new feature space that minimizes domain shift across datasets. The second method transfers the pre-trained CNN encoder from the first FDA method to the DistilHuBERT model. The third method builds upon the second approach; however, instead of transferring the CNN encoder, it fine-tunes DistilHuBERT using the target dataset. The last method fine-tunes both the CNN and transformer layers of DistilHuBERT for feature adaptation. Our emotion classifier incorporates attentive statistical pooling and Maxout (as a dimension reduction block), followed by a softmax layer. We implement CDA by updating the classifier by utilizing parts of the target dataset.Our datasets comprise EMODB, IEMOCAP, and ShEMO. Our fourth proposed FDA method, when combined with CDA, achieves the highest accuracy among our methods. It reaches 92.01 % on the EMODB dataset when using ShEMO as the source.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
Related Concept Videos
Labeling Emotion
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Systems-II
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as: