Related Experiment Videos
Speech Emotion Recognition Using Hybrid VMD and EWT Based Cepstral Feature Extraction
Siba Prasad Mishra1, Pankaj Warule2, Sudhansu Sekhar Nayak3
1School of Electronics Engineering (SENSE), Vellore Institute of Technology, Amaravathi, Andhra Pradesh, India.
Objective:
Speech Emotion Recognition (SER) has gained significant research attention over the past three decades owing to its diverse real-world applications, including human-computer interaction, healthcare, call centers, automotive systems, education, and security. The primary goal of SER is to accurately identify human emotions from speech and enhance emotion classification performance.
Study Design:
Recent studies have shown that signal decomposition-based feature extraction methods are more effective at capturing emotional cues than direct speech-level feature extraction, as different decomposition techniques can uncover complementary emotional information across the various frequency components of speech.
Methods:
Motivated by this, the present study integrates two signal decomposition approaches, such as variational mode decomposition (VMD) and empirical wavelet transform (EWT), to decompose the speech signal frame (SSF) into multiple sub-signals, referred to as modes or intrinsic mode functions (IMFs). From these decomposed signals, features such as mel frequency cepstral coefficients (MFCC) and mel frequency magnitude coefficient (MFMC) are extracted. The VMD- and EWT-derived features are then combined to form the proposed variational mode empirical wavelet transform-based mel frequency cepstral and magnitude coefficient (PVEWMFCMC) features, which are utilized for emotion classification using a Deep Neural Network (DNN) classifier.
Results:
Experimental evaluations demonstrate that the proposed PVEWMFCMC features, combined with a DNN model, achieved speaker-dependent classification accuracies of 91.76%, 86.92%, and 81.53% on the EMO-DB, EMOVO, and RAVDESS datasets, respectively. To further evaluate the generalization capability of the proposed approach, speaker-independent evaluation using the Leave-One-Speaker-Out (LOSO) protocol was also performed, achieving accuracies of 74.59%, 58.43%, and 56.79% on the EMO-DB, EMOVO, and RAVDESS datasets, respectively.