Related Experiment Video
Updated: Sep 12, 2025

08:22
Author Spotlight: Advancing the Study of Brain-Heart Interplay with a Comprehensive EEGLAB Plugin for Multimodal Signal Analysis
Published on: April 26, 2024
2.2K
Evaluating crowdsourcing for ICU EEG annotation: A comparison with expert performance
Wan-Yee Kong1,2, Fábio A Nascimento3, Aaron Struck4
1Beth Israel Deaconess Medical Center, Boston, Massachusetts, USA.
Epilepsia
|August 6, 2025
Summary
Crowdsourcing EEG annotations using a mobile app showed that weighted majority votes from non-experts were comparable to expert performance in identifying seizures and rhythmic patterns. This approach could accelerate the creation of large datasets for automated detection algorithms.
Area of Science:
- Neuroscience
- Medical Informatics
- Computational Biology
Background:
- Accurate detection of seizures and rhythmic or periodic patterns (SRPPs) on electroencephalography (EEG) is vital for managing critically ill neurological patients.
- Automated EEG analysis methods require large, expert-annotated datasets, but neurophysiologist availability limits expert annotation.
- Crowdsourcing offers a potential solution to scale up EEG data annotation.
Purpose of the Study:
- To evaluate the feasibility of using crowdsourcing for annotating EEG recordings.
- To compare the performance of non-expert crowdsourced annotations against expert neurophysiologists in identifying six types of SRPPs.
Main Methods:
- An EEG scoring contest was conducted via a mobile app, engaging 1542 participants (8 experts, 1534 non-experts).
- Participants annotated 478,834 short EEG epochs across six SRPPs: seizures, generalized and lateralized periodic discharges (GPDs, LPDs), and generalized and lateralized rhythmic delta activity (GRDA, LRDA), plus 'Other'.
- Performance was assessed using pairwise agreement, Fleiss' kappa for experts, and accuracy comparisons between experts and the crowd via individual and weighted majority votes.
Main Results:
- The crowd's individual, non-weighted votes were inferior to experts for overall and specific SRPP identification.
- Using weighted majority votes, the crowd achieved non-inferior overall SRPP identification accuracy (.70, 95% CI: .69-.70) compared to experts (.68, 95% CI: .68-.70).
- The crowd matched or exceeded expert performance for most SRPPs, excluding LPDs and 'Other'; no single expert outperformed the crowd overall.
Conclusions:
- Crowd reviewers show promise for achieving expert-level EEG annotations, potentially enabling the development of larger, more diverse datasets for automated detection algorithms.
- This proof-of-concept study suggests crowdsourcing is a viable method for EEG annotation.
- Further research is needed to address challenges like participant calibration and the absence of gold-standard labels in real-world applications.

