Related Experiment Video
Updated: Feb 14, 2026

The Mouse Stroke Unit Protocol with Standardized Neurological Scoring for Translational Mouse Stroke Studies
Published on: February 7, 2025
Towards phenotyping stroke: Leveraging data from a large-scale epidemiological study to detect stroke diagnosis
Yizhao Ni1,2, Kathleen Alwell3, Charles J Moomaw3
1Department of Biomedical Informatics, Cincinnati Children's Hospital Medical Center, Cincinnati, Ohio, United States of America.
Machine learning accurately detects stroke cases and subtypes from hospitalization data, outperforming traditional methods. This approach enhances stroke diagnosis for future genetic research.
Area of Science:
- Medical informatics
- Epidemiology
- Machine learning in healthcare
Background:
- Stroke diagnosis relies on accurate phenotyping for effective treatment and research.
- Traditional methods like ICD-9 codes have limitations in precision.
- Large-scale epidemiological studies generate extensive data valuable for refining diagnostic tools.
Purpose of the Study:
- To develop and assess machine learning (ML) algorithms for detecting stroke cases and subtypes using hospitalization data.
- To evaluate ML algorithm performance against established methods (ICD-9 codes, study nurses) using real-world epidemiological data.
- To identify key predictors for developing high-precision stroke phenotypic signatures.
Main Methods:
- Utilized 8,131 hospitalization events from the Greater Cincinnati/Northern Kentucky Stroke Study (2005, 2010).
- Extracted demographic and clinical variables from patient medical records.
- Trained ML algorithms to predict stroke cases and subtypes, validated against physician-adjudicated labels.
Main Results:
- The best ML algorithm achieved 88.57% accuracy for stroke case detection and 87.39% for subtype detection.
- ML algorithms significantly outperformed ICD-9 code classifications (P<0.001) across all measures.
- ML performance was comparable to study nurses, offering a better precision-recall tradeoff.
Conclusions:
- Machine learning demonstrates significant promise for improving stroke diagnosis accuracy from hospitalization data.
- This approach can enhance statistical power for subsequent genetic and genomic studies.
- Identified predictive variables can guide the development of precise stroke phenotyping algorithms.
Related Concept Videos
Regulation of Stroke Volume
Preload refers to the degree of stretch on the heart before it contracts. It's analogous to the stretching of a rubber band; the more it's stretched, the more forcefully it snaps back. This concept is encapsulated in the Frank-Starling law of the...
Statistical Methods for Analyzing Epidemiological Data
Cardiac Output and Stroke Volume
In an average resting adult male, the typical cardiac...
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Confounding in Epidemiological Studies
Bias in Epidemiological Studies

