Related Experiment Video
Updated: Jun 29, 2026

Motor Imagery Performance Through Embodied Digital Twins in a Virtual Reality-Enabled Brain-Computer Interface Environment
Published on: May 10, 2024
A Robust Multi-Branch CNN-LSTM Architecture for Cross-Subject Motor Imagery Classification
Simone Zini1, Federico Bidone1, Paolo Napoletano1
1Department of Informatics, Systems and Communication, University of Milano-Bicocca, Viale Sarca 336, 20126 Milano, Italy.
None:
Brain-computer interfaces (BCIs) based on motor imagery (MI) aim to convert electroencephalographic (EEG) activity into reliable device commands across users and recording setups. However, low signal-to-noise ratio and strong inter-subject variability still limit true "plug-and-play" deployment without lengthy calibration. To address these challenges, we propose a multi-branch convolutional long short-term memory (CNN-LSTM) architecture that jointly performs multi-scale temporal feature extraction and within-trial sequence modeling. The model employs four parallel 1D convolutional branches with distinct kernel sizes, each followed by an LSTM module and late fusion, combined with group normalization and supervision over sequences of sub-windows within each trial. We evaluate the approach on the EEG Motor Movement/Imagery (EEGMMI) dataset from PhysioNet under strictly subject-independent conditions, and on the ISLab-MI Dataset, a 32-channel wearable-EEG collection designed to assess cross-setup robustness. On EEGMMI, the network achieves up to 82.63% accuracy for binary left/right MI and 74.10% for a four-class task using 4 s trials under 5-fold cross-validation, outperforming an EEGNet-style baseline by 1-10% depending on class count and window length. Under a leave-one-subject-out protocol, the model attains 74.9% mean accuracy for a three-class MI task. Zero-shot transfer to ISLab-MI yields 64.60% and 63.02% accuracy in three- and four-class settings, respectively, while brief subject-specific fine-tuning using only 20% of each session improves performance to 81.38% and 73.48%. These findings show that combining multi-scale convolutional feature extraction with explicit sequence modeling and robust normalization yields accurate, data-efficient, and portable MI decoders suitable for practical BCI applications.
