Related Experiment Video
Updated: Aug 27, 2026

Electromagnetic Source Imaging in Presurgical Evaluation of Children with Drug-Resistant Epilepsy
Published on: September 20, 2024
Assessing the performance of artificial intelligence in detecting electrographic status epilepticus when human
Khalid Alsherbini1, Michelle Armenta Salas2, Suganya Karunakaran2
1Department of Neurology, Banner Health, University of Arizona, Phoenix, AZ, United States.
Objective:
Generating reference standards to train and evaluate the accuracy of artificial intelligence (AI) algorithms poses a significant challenge. Particularly when interpreting complex signals like electroencephalography (EEG), where interrater variability is considerable. We aimed to characterize the impact of interrater variability when evaluating the performance of an AI algorithm for detecting electrographic status epilepticus in point-of-care (POC) limited-montage EEG.
Methods:
We analyzed 604 EEGs collected using a POC EEG system (Ceribell Inc.). Each EEG was independently reviewed by 5-7 blinded experts, who annotated seizures and completed standardized assessments. The EEGs were later analyzed by the Clarity AI algorithm (version 7), trained on separate EEG dataset. Interrater agreement was assessed using Gwet's AC1. Sensitivity and specificity for detecting electrographic status epilepticus (ESE) were estimated for the reviewers and the AI algorithm. Multiple evaluation schemes were employed, including simple majority consensus of the full group (group majority), simple majority using leave-one-out analysis, and 2-of-3 majority with bootstrap sampling.
Results:
Of 604 POC EEGs, the group majority identified 8 cases (1.3%) meeting ACNS criteria for ESE and 14 (2.3%) with seizures, the rest were classified as normal/slowing (80.1%), highly epileptiform patterns (6.6%), or other findings (4.3%). Twentynine cases lacked consensus interpretation. Interrater agreement among reviewers was 0.67-0.68. Compared to the group majority, the AI algorithm showed higher sensitivity (median 100%) with lower specificity (93.5%) than individual reviewers (median sensitivity 60%, specificity 98.7%). Using all possible 3-reviewer combinations to define majority agreement, the number of EEGs identified as ESE varied widely (4-18 cases, median = 10). The AI algorithm consistently achieved significantly higher sensitivity than external human reviewers (71.4% vs. 50%, p < 0.001). However, the AI's specificity, while still high (median = 93.9%), was slightly lower than that of human reviewers (median = 98.4%, p < 0.001), though the AI's specificity had consistently narrower spread.
Discussion:
This study highlights the challenges of defining the correct answer for EEG interpretation, especially for ESE, and its consequences for evaluating seizure detection AI tools. Future work should explore how AI assistance impacts human interpretation, particularly in reducing interrater variability across a broader range of EEG patterns.

