Related Experiment Videos
Dual-Branch Cross-Attention Network for Electroencephalography and Facial Multimodal Emotion Recognition During
Lei Zha1, YuXue Feng2, Zhe Liang3
1College of Art, Minzu Normal University of Xingyi; Department of Electrical, Electronic and Automatic Engineering, University of Girona; u6107232@campus.udg.edu.
None:
Emotion recognition during painting viewing remains difficult because aesthetic responses are dynamic, context-dependent, and only partly observable through outward behavior. Unimodal approaches based only on facial expressions or only on physiological signals, therefore, often provide incomplete representations of the viewer's experience. This article describes a reproducible method for constructing a dual-branch cross-attention network that integrates facial video with electroencephalography (EEG) data for multimodal emotion recognition in a painting exhibition context. The protocol begins with the establishment of a synchronized acquisition environment using a 64-channel EEG system and a high-resolution camera for concurrent recording during painting viewing. The subsequent procedure details signal preprocessing, including EEG filtering, segmentation, and differential entropy feature extraction, together with facial detection, alignment, and frame preparation for video analysis. The neural network contains two modality-specific branches. The EEG branch uses a convolutional neural network and a bidirectional long short-term memory network to model spatial-temporal neural features, whereas the facial branch uses a coordinate attention-enhanced lightweight convolutional network to extract visual expression features. A multi-head cross-attention module is then used to fuse the two streams by learning intermodal relevance. Under the experimental conditions reported here, the multimodal framework achieved higher classification accuracy than the unimodal comparison models in four-category emotion recognition. This method is most suitable for offline analysis in controlled acquisition settings and provides a practical framework for affective computing, empirical aesthetics, and exhibition-oriented human-computer interaction research.