Related Experiment Video
Updated: Jul 1, 2026

Artificial Intelligence-Based System for Detecting Attention Levels in Students
Published on: December 15, 2023
A two-stage deep learning approach for facial emotion recognition in RAVDESS videos
Achraf Jallaglag1, My Abdelouahed Sabri2, Ali Yahyaouy3
1Computer Science Department, Faculty of Sciences, University Sidi Mohamed Ben Abdallah, Fez, Morocco. Achraf.jallaglag@usmba.ac.ma.
Abstract:
Video-based emotion recognition is an important topic in affective computing, with applications in human-computer interaction, mental health, and multimedia systems. In this work, we propose a two-stage deep learning approach that extracts spatial features from individual frames and leverages temporal consistency across video sequences for facial expression recognition using the RAVDESS dataset. First, videos are split into frames, and a fine-tuned VGG16 CNN extracts discriminative spatial features from each frame. Second, these features are aggregated into sequences and the frame-level predictions are aggregated at the video level using a majority voting strategy to ensure temporal consistency. Our experiments show that the proposed method achieves 93.6% accuracy at the frame level and 98.1% at the video level, outperforming baseline models and remaining competitive with recent state-of-the-art approaches. Temporal aggregation helps reduce misclassifications of subtle emotions, while fine-tuning improves feature extraction. The approach is computationally efficient and provides a solid foundation for future research in multimodal emotion recognition and advanced video-level aggregation.
Related Concept Videos
Facial Feedback Hypothesis
Labeling Emotion