Related Experiment Video
Updated: Oct 12, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Multimodal Emotion Recognition on RAVDESS Dataset Using Transfer Learning.
Cristina Luna-Jiménez1, David Griol2, Zoraida Callejas2
1Grupo de Tecnología del Habla y Aprendizaje Automático (THAU Group), Information Processing and Telecommunications Center, E.T.S.I. de Telecomunicación, Universidad Politécnica de Madrid, Avda. Complutense 30, 28040 Madrid, Spain.
This study introduces a multimodal emotion recognition system using speech and facial data. Combining these modalities achieved 80.08% accuracy, enhancing emotion detection for applications in healthcare and road safety.
Area of Science:
- Computer Science
- Artificial Intelligence
- Human-Computer Interaction
Background:
- Emotion recognition is crucial for applications in healthcare and road safety.
- Multimodal approaches combining speech and facial data offer potential for improved accuracy.
Purpose of the Study:
- To develop and evaluate a multimodal emotion recognition system using speech and facial information.
- To investigate the effectiveness of transfer learning techniques for speech emotion recognition.
- To propose a novel framework for facial emotion recognition and combine it with speech for enhanced performance.
Main Methods:
- Speech emotion recognition: Evaluated transfer learning (embedding extraction, Fine-Tuning) using CNN-14 (PANNs framework).
- Facial emotion recognition: Proposed a framework with a pre-trained Spatial Transformer Network and a bi-LSTM with attention.
- Multimodal fusion: Employed a late fusion strategy to combine speech and facial modalities.
Main Results:
- Fine-tuning CNN-14 yielded the most robust speech emotion recognition results.
- The facial emotion recognition framework showed promise, though frame-based systems require further research for video tasks.
- The combined multimodal system achieved 80.08% accuracy on the RAVDESS dataset for classifying eight emotions using subject-wise 5-fold cross-validation.
Conclusions:
- Speech and facial modalities contain significant information for detecting emotional states.
- Combining speech and facial recognition through late fusion effectively improves overall system performance.
- The findings highlight the potential of multimodal emotion recognition in real-world applications.
Related Concept Videos
Labeling Emotion
Physiology of Emotion
Autonomic Nervous System
The autonomic nervous system (ANS) plays a critical role in emotional responses by regulating involuntary physiological functions. It consists of two main components: the sympathetic and parasympathetic systems. The sympathetic system...
Multi-input and Multi-variable systems
In the absence...
Introduction to Motivation and Emotion
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Cognitive Theories: Lazarus Mediational Theory of Emotion
Cognitive Appraisal and Emotional Response
Lazarus proposed that...

