Related Experiment Video
Updated: Nov 24, 2025

06:37
Electroencephalography Network Indices as Biomarkers of Upper Limb Impairment in Chronic Stroke
Published on: July 14, 2023
1.1K
Comparative analysis, applications, and interpretation of electronic health record-based stroke phenotyping methods
Phyllis M Thangaraj1,2, Benjamin R Kummer3, Tal Lorberbaum1,2
1Department of Biomedical Informatics, Columbia University, 622 W 168th St., PH-20, New York, NY, 10032, USA.
Biodata Mining
|December 29, 2020
Summary
Machine learning accurately identifies acute ischemic stroke (AIS) patients using electronic health records. This automated method is generalizable and efficient for clinical research cohort identification.
Area of Science:
- Biomedical informatics
- Clinical research methodology
- Artificial intelligence in healthcare
Background:
- Accurate identification of acute ischemic stroke (AIS) patient cohorts is critical for clinical investigations.
- Current phenotyping methods are laborious and lack generalizability.
- Electronic health records (EHRs) offer a new approach for automated cohort identification.
Purpose of the Study:
- To systematically compare machine learning algorithms and case-control combinations for phenotyping AIS patients using EHR data.
- To evaluate the generalizability and accuracy of automated phenotyping methods.
- To assess the utility of diagnosis codes in training classifier models.
Main Methods:
- Developed and evaluated machine learning models using structured EHR data from a tertiary-care hospital system.
- Tested 75 different case-control and classifier combinations for AIS patient identification.
- Externally validated models using UK Biobank data to detect AIS patients without diagnosis codes.
Main Results:
- Machine learning models achieved a mean AUROC of 0.963 and average precision of 0.790 with minimal feature processing.
- Classifiers trained with AIS diagnosis codes and controls without cerebrovascular disease codes yielded the best F1 score (0.832).
- External validation showed significant enrichment (60-150 fold) of AIS patients without diagnosis codes in top-predicted cohorts.
Conclusions:
- Machine learning algorithms provide a generalizable and accurate method for identifying AIS patients without manual feature curation.
- Diagnosis codes can be effectively used to train classifier models when a specific AIS patient set is unavailable.

