Related Experiment Video
Updated: Aug 28, 2026

Motor Imagery Performance Through Embodied Digital Twins in a Virtual Reality-Enabled Brain-Computer Interface Environment
Published on: May 10, 2024
Evaluating Validation Strategies in Motor Imagery EEG: A Full-Cohort GAF-PLV Analysis and Matched Sensitivity Study
Wenwen Chang1, Hesam Akbari2, Muhammad Tariq Sadiq3
1School of Electronic and Information Engineering, Lanzhou Jiaotong University, Lanzhou 730070, China.
Abstract:
Background: Performance estimates in motor-imagery electroencephalography (MI-EEG) can depend strongly on how observations are partitioned for training, model selection, and testing. Random sample- or window-level splitting may place data from the same participant in different folds and therefore does not answer the same question as evaluation on previously unseen participants. Methods: We evaluated a previously developed Gramian angular field-phase-locking value (GAF-PLV) classifier on the retained full cohort (N=105) using binary left-versus-right MI and leave-one-subject-out cross-validation (LOSO). Separately, a predefined, outcome-independent subset (N=30) was used for a matched sensitivity analysis of eight classifiers, binary and four-class tasks, and three validation strategies: random five-fold cross-validation, LOSO, and nested LOSO with subject-grouped inner model selection. Results: In the full-cohort GAF-PLV analysis, mean accuracy was 58.07% ± 8.27% and Macro-F1 was 53.48% ± 11.19%, with substantial between-subject variability. In the predefined matched subset, performance estimates and numerical model rankings changed across validation strategies. For example, the numerically highest binary-accuracy model was ShallowConvNet under random five-fold cross-validation, DeepConvNet under LOSO, and ATCNet under nested LOSO. Conclusions: Random within-cohort classification and generalisation to previously unseen subjects are distinct evaluation targets. MI-EEG reports should state the cohort, partition unit, validation design, and model-selection procedure. Rankings in the multi-model analysis are conditional on the predefined 30-subject subset and common 0.4 s input setting and are not presented as definitive full-cohort or architecture-optimal rankings.
