Related Experiment Video
Updated: Oct 8, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
AttendAffectNet-Emotion Prediction of Movie Viewers Using Multimodal Fusion with Self-Attention.
Ha Thi Phuong Thao1, B T Balamurali2, Gemma Roig3
1Information Systems Technology and Design, Singapore University of Technology and Design, 8 Somapah Rd, Singapore 48737, Singapore.
This study introduces AttendAffectNet (AAN), a novel neural network for predicting viewer emotions from movies. Audio features are most influential, and combining multimodal inputs improves prediction accuracy, with Feature AAN showing superior performance.
Area of Science:
- Affective computing
- Multimodal machine learning
- Computational psychology
Background:
- Predicting viewer affective responses from movie content is challenging.
- Existing methods often overlook correlations within and between temporal and multimodal inputs.
- Developing robust models requires integrating diverse features like visual, audio, and text.
Purpose of the Study:
- To propose and evaluate AttendAffectNet (AAN), a novel neural network architecture for predicting viewer emotions from movie content.
- To explore the impact of self-attention mechanisms on multimodal and temporal feature correlations.
- To compare the performance of different AAN variants (Feature AAN, Temporal AAN, Mixed AAN) and identify key predictive features.
Main Methods:
- Developed AttendAffectNet (AAN), a neural network utilizing self-attention for emotion prediction.
- Extracted visual, audio, and text features from movie data.
- Implemented and compared three AAN variants: Feature AAN, Temporal AAN, and Mixed AAN.
- Validated models on the MediaEval 2016 and COGNIMUSE datasets.
Main Results:
- Audio features demonstrated greater influence on emotion prediction than visual or text features.
- Multimodal models integrating all feature types outperformed unimodal models.
- The Feature AAN variant achieved superior performance across datasets, outperforming other AAN variants and baseline models.
- Feature AAN showed particular strength in predicting the valence dimension.
Conclusions:
- The proposed AttendAffectNet (AAN) effectively predicts viewer affective responses using multimodal movie content.
- Integrating features from different modalities and considering their interrelationships is crucial for accurate emotion prediction.
- Audio modality plays a significant role, and the Feature AAN architecture offers a promising approach for multimodal emotion recognition.
More Related Videos
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Related Concept Videos
The Influence of Cognition on Affect
Labeling Emotion
Facial Feedback Hypothesis
The Influence of Affect on Cognition
Role of Affect in Interpersonal Attraction
Cognitive Theories: Schachter-Singer Theory of Emotion
Physiological Arousal and Cognitive Labeling
According to this theory, when an individual experiences...