Related Experiment Video
Updated: Jan 15, 2026

Using Electroencephalography Measurements and High-quality Video Recording for Analyzing Visual Perception of Media Content
Published on: May 26, 2018
Detecting mind wandering via EEG and facial video features
Shaohua Tang1,2,3, Chunbo Jiang4, Zheng Li5,6
1Department of Systems Science, Faculty of Arts and Sciences, Beijing Normal University, Zhuhai, 519087, Guangdong, China.
Purpose:
Mind wandering (MW), a common cognitive phenomenon marked by a shift of attention away from the task at hand, poses significant challenges in online educational settings. This study aims to advance MW detection by developing a classification scheme that leverages multimodal data, including electroencephalograph (EEG) signals and facial video recorded using a commercial off-the-shelf webcam. Additionally, this study provides an in-depth analysis of feature contributions and explores the correlation between self-reported introspective confidence, mental state stability, and classification performance, offering deeper insights into MW detection.
Methods:
Data were collected from 26 college students during a video-based learning task, interspersed with modified experience sampling probes. To enhance the sample size and address autocorrelation in EEG signals, a probe-based sample extraction method was applied. MW classification was performed using a random forest algorithm, with features derived from both EEG signals and facial video recordings. Model performance was evaluated using within-participant tenfold cross-validation and leave-one-participant-out (LOPO) cross-validation.
Results:
The combination of EEG and video features yielded better performance (AUC = 0.68 for within-participant; AUC = 0.56 for LOPO) compared to using EEG or video alone. Individual differences significantly influenced performance, with a 10% increase in AUC observed when training data included samples from the evaluated individual in augmented LOPO cross-validation. Introspective confidence levels positively correlated with classification performance, while mental state temporal stability was associated with improved cross-participant performance. Additionally, the size of the training set positively correlated with cross-participant performance when combining EEG and video features.
Conclusion:
These findings underscore the potential of multimodal approaches for MW detection and highlight the importance of individual differences and data diversity in classifier training. The study provides actionable insights into improving MW detection systems for real-world applications in educational settings.

