Related Experiment Video
Updated: Aug 18, 2026

Simultaneous Scalp Electroencephalography (EEG), Electromyography (EMG), and Whole-body Segmental Inertial Recording for Multi-modal Neural Decoding
Published on: July 26, 2013
Application of a parametric supermatrix multi-task multimodal deep learning model based on electroencephalogram and
Xue Li1,2, Piqiang Gong1, Chuantao Li3
1Medical Security Center, The 940th Hospital of Joint Logistics Support Force of Chinese People's Liberation Army, Lanzhou, Gansu, China.
Abstract:
Accurate emotion recognition under limited computational resources remains challenging in applications such as healthcare, human-computer interaction, and intelligent systems. To address this issue, we propose a parameterized supermatrix multi-task multimodal emotion recognition model (PH-MTM) that fuses electroencephalography (EEG) and peripheral physiological signals (PPS). Physiological signals from the DEAP dataset are preprocessed to extract differential entropy (DE) and power spectral density (PSD) features, which are organized into a "frequency band number × 8 × 9" grid according to the international 10-20 system. An improved squeeze-and-excitation (SE) attention mechanism adaptively assigns weights to different frequency bands, while a convolutional neural network (CNN) performs deep fusion of cross-band and cross-channel information. To reduce model complexity without sacrificing performance, a lightweight fully connected layer based on matrix decomposition is integrated into a multi-task learning framework. Because DEAP includes only 32-channel EEG signals, additional experiments are conducted on the 62-channel SEED dataset to evaluate the scalability and robustness of the frequency-band fusion strategy. PH-MTM achieves 95.99 and 96.51% accuracy for valence and arousal classification on DEAP, and 94.25% accuracy for three-class emotion recognition on SEED. Fusion analysis of EEG with different PPS types shows that EEG combined with electromyography (EMG) performs best, reaching 99.69 and 99.74% accuracy for valence and arousal, respectively. Compared with existing methods, the proposed model demonstrates improved recognition accuracy while maintaining low computational cost. In addition, SHAP visualization is applied to analyze the contribution of different frequency-band features, offering further insight into multimodal physiological signal-based emotion recognition.
