Related Experiment Video
Updated: Sep 16, 2026

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
KTU-MEDAFE: A Newly Developed Multimodal Dataset for Emotion Recognition Using EEG-Speech Decision-Level Fusion
Bahar Hatipoglu Yilmaz1, Betul Mumcu2, Busra Ozkellekci1
1Department of Computer Engineering, Engineering Faculty, Karadeniz Technical University, 61080 Trabzon, Türkiye.
Abstract:
One of the central challenges in affective computing is achieving reliable emotion recognition for natural and effective human-computer interaction. In this study, we introduce KTU-MEDAFE (Karadeniz Technical University Multimodal Emotion Dataset using Audio, Facial Images, and EEG), a newly developed multimodal dataset containing synchronized EEG signals, speech recordings, and facial videos collected from 40 participants under controlled emotional elicitation conditions. The dataset includes Turkish emotional speech and two recording sessions conducted on separate days, providing a language-specific resource that supports both participant-dependent baseline evaluation and future session-separated analysis. Although KTU-MEDAFE comprises three modalities, the present study focuses on EEG and speech integration. EEG and speech recordings meeting signal quality criteria were transformed into image representations using the Angle-Amplitude Graph (AAG) method and classified using transfer learning with ResNet-50 and GoogLeNet architectures. To exploit complementary information across modalities, multiple decision-level fusion strategies were evaluated. Experimental findings show that multimodal fusion provides higher average classification performance than unimodal EEG and speech models across the evaluated binary emotion pairs, with performance varying according to subject, fusion strategy, and model architecture. Overall, the results support the potential benefit of combining EEG and speech for multimodal emotion recognition while highlighting substantial subject-dependent variability in classification performance.

