Related Experiment Video
Updated: May 5, 2026

05:51
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
9.0K
A Combined CNN Architecture for Speech Emotion Recognition.
Rolinson Begazo1, Ana Aguilera2,3, Irvin Dongo1,4
1Electrical and Electronics Engineering Department, Universidad Católica San Pablo, Arequipa 04001, Peru.
Sensors (Basel, Switzerland)
|September 14, 2024
Summary
This study enhances emotion recognition in speech using deep learning. A novel approach combining spectral features and spectrogram images achieved 96% accuracy, improving human-computer interaction.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Human-Computer Interaction
Background:
- Emotion recognition from speech is crucial for Human-Computer Interaction (HCI).
- Existing deep learning methods face challenges with data quantity, diversity, and feature selection standards.
- Designing effective neural network architectures for speech emotion recognition remains complex.
Purpose of the Study:
- To address limitations in current speech emotion recognition techniques using deep learning.
- To propose a comprehensive approach involving data preprocessing, feature selection, and a novel neural network architecture.
- To develop and utilize a unified dataset (EmoDSc) for robust emotion recognition.
Main Methods:
- Construction of a unified dataset (EmoDSc) by combining existing speech emotion databases.
- Investigation of the synergy between spectral features and spectrogram images for emotion recognition.
- Development of a hybrid neural network architecture integrating 1D Convolutional Neural Network (CNN1D), 2D CNN (CNN2D), and Multilayer Perceptron (MLP) to fuse spectral and image-based features.
Main Results:
- Individual analysis showed weighted accuracies of 89% for spectral features and 90% for spectrogram images.
- The proposed hybrid model, utilizing the EmoDSc dataset, achieved a remarkable weighted accuracy of 96%.
- The fused approach significantly outperformed models relying on isolated feature types.
Conclusions:
- The proposed deep learning approach, combining spectral features and spectrogram images via a hybrid CNN-MLP architecture, significantly advances speech emotion recognition.
- The unified EmoDSc dataset provides a valuable resource for training and evaluating speech emotion recognition models.
- This study offers a robust solution to improve the accuracy and reliability of emotion recognition in human-computer interaction systems.

