Related Experiment Video
Updated: Dec 23, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Speech Emotion Recognition Based on Selective Interpolation Synthetic Minority Over-Sampling Technique in Small
Zhen-Tao Liu1,2, Bao-Han Wu1,2, Dan-Yun Li1,2
1School of Automation, China University of Geosciences, Wuhan 430074, China.
This study introduces a novel speech emotion recognition model for small datasets. It effectively handles data imbalance and redundant features, achieving high accuracy across multiple databases.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Machine Learning
Background:
- Speech emotion recognition (SER) faces challenges with imbalanced data and feature redundancy across diverse applications.
- Existing SER models are often tailored to specific sample conditions, limiting generalizability.
- A robust SER model for small sample environments is needed.
Purpose of the Study:
- To propose an effective speech emotion recognition model for small sample environments.
- To address data imbalance and feature redundancy in SER.
- To improve the accuracy and robustness of SER systems.
Main Methods:
- Developed a data imbalance processing method using Selective Interpolation Synthetic Minority Over-sampling Technique (SISMOTE).
- Implemented a feature selection method combining variance analysis and Gradient Boosting Decision Tree (GBDT) to remove redundant features.
- Evaluated the proposed model on CASIA, Emo-DB, and SAVEE speech emotion databases.
Main Results:
- Achieved high speaker-dependent speech emotion recognition accuracies: 90.28% on CASIA, 75.00% on SAVEE, and 85.82% on Emo-DB.
- Demonstrated superior performance compared to several state-of-the-art methods.
- Effectively reduced the impact of sample imbalance and excluded poorly representative features.
Conclusions:
- The proposed SISMOTE and GBDT-based feature selection method significantly enhances SER performance in small sample scenarios.
- The model offers a robust and accurate solution for emotion recognition from speech, outperforming existing approaches.
- This work provides a valuable contribution to the field of affective computing and human-computer interaction.
More Related Videos
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Reconstruction of Signal using Interpolation
Sampling Methods: Overview
In analytical chemistry, the choice of...
Upsampling
Sampling Theorem
Labeling Emotion