Related Experiment Video
Updated: Oct 27, 2025

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
Robust Multimodal Emotion Recognition from Conversation with Transformer-Based Crossmodality Fusion
Baijun Xie1, Mariia Sidulova1, Chung Hyuk Park1
1Department of Biomedical Engineering, School of Engineering and Applied Science, George Washington University, Washington, DC 20052, USA.
Abstract:
Decades of scientific research have been conducted on developing and evaluating methods for automated emotion recognition. With exponentially growing technology, there is a wide range of emerging applications that require emotional state recognition of the user. This paper investigates a robust approach for multimodal emotion recognition during a conversation. Three separate models for audio, video and text modalities are structured and fine-tuned on the MELD. In this paper, a transformer-based crossmodality fusion with the EmbraceNet architecture is employed to estimate the emotion. The proposed multimodal network architecture can achieve up to 65% accuracy, which significantly surpasses any of the unimodal models. We provide multiple evaluation techniques applied to our work to show that our model is robust and can even outperform the state-of-the-art models on the MELD.
Related Concept Videos
Labeling Emotion
Transformers
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
Cognitive Theories: Lazarus Mediational Theory of Emotion
Cognitive Appraisal and Emotional Response
Lazarus proposed that...
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Facial Feedback Hypothesis
Empathy

