Related Experiment Video
Updated: Sep 2, 2026

Brain Source Imaging in Preclinical Rat Models of Focal Epilepsy using High-Resolution EEG Recordings
Published on: June 6, 2015
SEDAT: a hybrid tokenizer for large EEG models
Muhammad Zulkifal Aziz1, Yue Zhuo2, Binwen Huang3
1School of Automation, Northwestern Polytechnical University, Xi'an 710072, Shaanxi, People's Republic of China.
Abstract:
Objective.The fidelity of neural representations learned by large EEG foundation models depends on how raw brain signals are tokenized. Existing methods suffer from arbitrary temporal boundaries misaligned with neural state transitions, neglecting inter-channel spatial information, and fixed segmentation criteria that fail to generalize across heterogeneous EEG paradigms.Approach.This study proposes the squeeze-and-excitation (SE)-data-adaptive Gaussian average filtering (DAGAF) adaptive tokenizer (SEDAT), a hybrid framework integrating SE-based spatial aggregation, DAGAF-based signal decomposition, instantaneous-frequency-guided adaptive segmentation, and Fourier-domain resampling into a single computationally efficient pipeline. SEDAT is evaluated across 10 heterogeneous EEG datasets spanning motor imagery, mental imagery, P300, slow cortical potentials, sleep staging, epilepsy, emotions recognition and Alzheimer diagnosis paradigms, using four large foundation models: LaBraM, EEGFormer, EEGPT, and NeuroGPT. It is benchmarked against five competitive baselines: fixed-length windowing (FLW), context segmentation, linear predictive coding-based tokenization, TFM-Tokenizer, and source informed segmentation.Main results.SEDAT achieves classification improvements of up to 15.3% over FLW and 1.2%-4.6% over the second-best method, with all comparisons reachingafter Benjamini-Hochberg correction. Token quality analysis confirms substantially improved feature separability, with Silhouette scores of 0.81-0.85 versus 0.33-0.48 for rigid baselines.Significance.Withcomplexity, SEDAT explores new applications for SE and DAGAF as tokenizers and provides a physiologically grounded and computationally practical tokenization solution for large-scale EEG foundation models.

