An extended clinical EEG dataset with 15,300 automatically labelled recordings for pathology decoding
Ann-Kathrin Kiessner1, Robin T Schirrmeister2, Lukas A W Gemein3
1Neuromedical AI Lab, Department of Neurosurgery, Medical Center - University of Freiburg, Faculty of Medicine, University of Freiburg, Engelbergerstr. 21, 79106 Freiburg, Germany; BrainLinks-BrainTools, IMBIT (Institute for Machine-Brain Interfacing Technology), University of Freiburg, Georges-Köhler-Allee 201, 79110 Freiburg, Germany; Autonomous Intelligent Systems, Computer Science Department - University of Freiburg, Faculty of Engineering, University of Freiburg, Georges-Köhler-Allee 80, 79110 Freiburg, Germany.
Training machine learning models on larger, automatically labeled electroencephalogram (EEG) datasets improves diagnostic accuracy for EEG pathology. This study introduces a new, expanded dataset (TUABEXB) to enhance deep learning model generalization in EEG analysis.
Area of Science:
- Neuroscience
- Computer Science
- Medical Informatics
Background:
- Machine learning (ML) for automated clinical EEG analysis is a rapidly advancing field.
- Existing datasets like the Temple University Hospital Abnormal EEG Corpus (TUAB) are limited in size for robust ML model generalization.
- The impact of training on larger, automatically labeled EEG datasets on deep neural network performance remains underexplored.
Purpose of the Study:
- To create and evaluate a larger, automatically labeled dataset for EEG pathology decoding.
- To investigate the effect of training on this new dataset versus the existing TUAB dataset on deep convolutional neural network (ConvNet) performance.
- To provide an open-source dataset for advancing EEG machine learning research.
Main Methods:
- Automatically generated pathology labels for the Temple University Hospital EEG Corpus (TUEG) using a rule-based text classifier on medical reports.
- Created the TUH Abnormal Expansion EEG Corpus (TUABEX) with 15,300 automatically labeled recordings.
- Derived a balanced subset, TUABEXB (8,879 recordings), from TUABEX.
- Trained four established ConvNets on both TUAB and TUABEXB for binary EEG pathology classification.
Main Results:
- Training ConvNets on the automatically labeled TUABEXB dataset led to increased classification accuracies compared to training on the manually labeled TUAB dataset.
- Performance improvements were observed on both the TUABEXB and, for some architectures, the original TUAB dataset.
- The findings suggest that automatically labeling large datasets is an effective strategy for utilizing clinical EEG data.
Conclusions:
- Automatically labeled, large-scale EEG datasets can significantly enhance the generalization performance of deep learning models for pathology detection.
- The newly created TUABEXB dataset offers a valuable resource for the EEG machine learning community.
- This approach facilitates the efficient use of vast amounts of archived clinical EEG data for research and development.


