Related Experiment Video
Updated: Jan 13, 2026

05:51
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
9.4K
A Transformer-Based Multimodal Fusion Network for Emotion Recognition Using EEG and Facial Expressions in
Shuni Feng1, Qingzhou Wu2, Kailin Zhang1
1Tianjin Key Laboratory of Life and Health Detection, Life and Health Intelligent Research Institute, Tianjin University of Technology, Tianjin 300384, China.
Sensors (Basel, Switzerland)
|October 29, 2025
Summary
This study introduces a novel multimodal fusion neural network for emotion recognition in hearing-impaired individuals. The proposed method significantly improves accuracy by integrating EEG and facial expression data using an attention mechanism.
Area of Science:
- Neuroscience
- Artificial Intelligence
- Biomedical Engineering
Background:
- Hearing impairment presents unique challenges in emotional expression and perception.
- Single-modal emotion recognition methods are insufficient for complex environments.
- Developing effective emotion recognition for hearing-impaired individuals is crucial.
Purpose of the Study:
- To propose a multimodal fusion neural network for enhanced emotion recognition in hearing-impaired individuals.
- To leverage both electroencephalogram (EEG) and facial expression data.
- To improve upon existing emotion recognition techniques.
Main Methods:
- Developed a multimodal multi-head attention fusion neural network (MMHA-FNN).
- Utilized differential entropy (DE) and bilinear interpolation features as inputs.
- Employed an MBConv-based module for spatial-temporal brain region characteristics and a Transformer-based multi-head self-attention mechanism for cross-modal feature interaction.
Main Results:
- Achieved an average accuracy of 81.14% on a four-classification task (happy, sad, fear, calmness) using the MED-HI dataset.
- Significantly outperformed feature concatenation (71.02%) and decision layer fusion (69.45%).
- Demonstrated the effectiveness of attention-based feature layer interaction fusion.
Conclusions:
- EEG and facial expressions are complementary modalities for emotion recognition in hearing-impaired individuals.
- The MMHA-FNN effectively models cross-modal dependencies and enhances recognition performance.
- This approach offers a promising solution for improving emotional communication for the hearing impaired.

