Related Experiment Video
Updated: Nov 7, 2025

10:25
Brain Infarct Segmentation and Registration on MRI or CT for Lesion-symptom Mapping
Published on: September 25, 2019
48.7K
Discrepancies in Stroke Distribution and Dataset Origin in Machine Learning for Stroke.
Lohit Velagapudi1, Nikolaos Mouchtouris1, Michael P Baldassari1
1Department of Neurosurgery, Thomas Jefferson University, Philadelphia, PA.
Summary
Machine learning models for stroke lack generalizability due to biased training data. Future research must use diverse patient populations reflecting actual stroke distribution to improve clinical tool accuracy and reduce disparities.
Area of Science:
- Clinical research
- Machine learning
- Public health
Background:
- Machine learning (ML) algorithms require accurate, representative datasets for effective clinical application and broad generalizability.
- Disparities in healthcare access and outcomes, such as stroke prevalence, necessitate careful consideration in ML model development.
- Geographic and demographic representation in training data is crucial for equitable ML tool deployment.
Purpose of the Study:
- To review machine learning applications in stroke literature.
- To assess the geographic distribution of datasets and patient cohorts used in ML model training.
- To compare this distribution with stroke prevalence to identify potential disparities.
Main Methods:
- A systematic search of the PubMed database was conducted.
- Studies were screened based on titles and abstracts, followed by full-text assessment.
- 27 studies met the inclusion criteria after excluding those with non-US cohorts, reviews, or editorials.
Main Results:
- Included studies predominantly utilized patient data from specific US regions, with California (25.9%) and multicenter studies (22.2%) being most common.
- Geographic distribution of training data showed significant underrepresentation from states with higher stroke prevalence, such as Mississippi (4.3%).
- Conversely, areas with lower stroke prevalence, like California (2.6%), were overrepresented in the training datasets.
Conclusions:
- A significant disconnect exists between the geographic distribution of datasets used for training ML algorithms and the actual stroke prevalence across the US.
- This disparity can lead to biased ML tools with limited generalizability and accuracy.
- Future ML studies must incorporate diverse patient populations that mirror the unequal distribution of stroke risk factors to enhance clinical tool usability and ensure nationwide accuracy.

