Related Experiment Video
Updated: May 9, 2025

06:37
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
2.4K
Multi-modal sentiment recognition with residual gating network and emotion intensity attention
Yadi Wang1, Xiaoding Guo1, Xianhong Hou2
1School of Computer and Information Engineering, Henan University, Kaifeng, 475004, China; Henan Key Laboratory of Big Data Analysis and Processing, Kaifeng, 475004, China.
Summary
This study introduces MSRG, a novel multimodal emotion recognition framework. MSRG effectively processes complementary information and selects key features, significantly improving emotion prediction accuracy on benchmark datasets.
Area of Science:
- Artificial Intelligence
- Computer Science
- Affective Computing
Background:
- Multimodal emotion recognition leverages text, visual, and acoustic data.
- Existing methods struggle with integrating complementary modal information and managing long-term dependencies.
- Effective selection of joint modal features remains a challenge.
Purpose of the Study:
- To propose a new multimodal emotion recognition framework, MSRG.
- To address limitations in processing complementary information and selecting joint modal features.
- To enhance the accuracy of emotion prediction using fused multimodal data.
Main Methods:
- MSRG framework includes feature extraction (FE), emotional intensity attention (EIA), time-step level fusion (TLF), utterance level fusion (ULF), and sentiment inference module (SIM).
- EIA incorporates adaptive multimodal linear pooling (AMLP) and joint cross-attention fusion (JCAF) for dynamic feature weighting and fusion.
- TLF and ULF employ residual gating networks (RGN) and concatenation for time-step and utterance-level fusion, respectively.
Main Results:
- The proposed MSRG framework demonstrates superior prediction performance on the CMU-MOSI and CMU-MOSEI datasets.
- Adaptive multimodal fusion and joint cross-attention effectively capture inter-modal relationships.
- The hierarchical fusion strategy (time-step and utterance levels) enhances feature representation.
Conclusions:
- MSRG offers an effective solution for multimodal emotion recognition by improving information processing and feature selection.
- The framework's adaptive and cross-attention mechanisms are key to its enhanced performance.
- MSRG sets a new benchmark for emotion recognition accuracy on established multimodal datasets.
Keywords:
Emotion recognitionEmotional intensity attentionResidual gating networkUtterance level fusion
