Evaluating the clinical readiness of artificial intelligence in EEG-based epilepsy diagnosis
Jeet Bandhu Lahiri1, Puneet Agarwal2, Suman Kushwaha3
1School of Computing and Electrical Engineering, Indian Institute of Technology Mandi, Mandi, Himachal Pradesh, India.
Abstract:
Objective.Automated electroencephalography (EEG)-based epilepsy diagnosis has reported near-perfect accuracies for almost two decades on a benchmark dataset, yet virtually no system is used in routine care. We critically re-examined this translation gap by reproducing five widely cited Artificial Intelligence (AI) models spanning statistical feature extraction, classical machine-learning and deep-learning paradigms, and assessed their ability to generalise from the benchmark dataset to a newly curated, clinically verified scalp-EEG cohort.Approach.All models were implemented as originally described and trained with ten-fold cross-validation on the benchmark dataset. External validation was performed on our independently curated dataset made publicly available, comprising 30 subjects (15 epilepsy, 15 healthy) recorded with 19-channel scalp EEG under standard clinical protocols. We further examined the influence of subject-level data leakage by contrasting performance when training/testing samples overlapped with those when completely independent patient partitions were enforced. Accuracy, sensitivity, specificity, and AUC of ROC (with 95% Confidence Intervals) were the primary metrics.Main results.When transferred unchanged to the external cohort, overall accuracy fell from⩾94% on the benchmark dataset to 42%-53%, with sensitivities as low as 0.97% for the deep convolutional neural network and specificities dropping to 4% for time-frequency methods. Permitting subject overlap artificially elevated accuracy to 59%-96%, whereas strict patient separation reduced it to 41%-53% in 95% Confidence Interval. Deep-learning models exhibited the steepest decline, confirming over-fitting to subject-specific artefacts. Statistical feature-based approaches, though less affected, still under-performed clinically acceptable thresholds.Significance.Our results expose key translational barriers in AI for EEG-based epilepsy diagnosis-data leakage, acquisition bias, and overfitting to patient idiosyncrasies-leading to severe performance erosion on clinical data. Rigorous patient-independent validation, transparent reporting (aligned with CARE principles), and well-curated multi-channel scalp-EEG datasets are essential to ensure clinically dependable AI tools for epilepsy diagnosis.
More Related Videos
09:57Author Spotlight: Advancing Pediatric Epilepsy Surgery in Children Through Novel Biomarkers and Enhanced Localization
Published on: September 20, 2024
09:00Investigating the Function of Deep Cortical and Subcortical Structures Using Stereotactic Electroencephalography: Lessons from the Anterior Cingulate Cortex
Published on: April 15, 2015
