Development and validation of a stroke risk prediction model using regional healthcare big data and machine learning
Yunxia Duan1,2, Rui Wang1,3, Yumei Sun1
1School of Nursing, Peking University, Beijing, China.
Objectives:
This study aimed to develop and validate a stroke risk prediction model based on machine learning (ML) and regional healthcare big data, and determine whether it may improve the prediction performance compared with the conventional Logistic Regression (LR) model.
Methods:
This retrospective cohort study analyzed data from the CHinese Electronic health Records Research in Yinzhou (CHERRY) (2015-2021). We included adults aged 18-75 from the platform who had established records before 2015. Individuals with pre-existing stroke, key data absence, or excessive missingness (>30 %) were excluded. Data on demographic, clinical measures, lifestyle factors, comorbidities, and family history of stroke were collected. Variable selection was performed in two stages: an initial screening via univariate analysis, followed by a prioritization of variables based on clinical relevance and actionability, with a focus on those that are modifiable. Stroke prediction models were developed using LR and four ML algorithms: Decision Tree (DT), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Back Propagation Neural Network (BPNN). The dataset was split 7:3 for training and validation sets. Performance was assessed using receiver operating characteristic (ROC) curves, calibration, and confusion matrices, and the cutoff value was determined by Youden's index to classify risk groups.
Results:
The study cohort comprised 92,172 participants with 436 incident stroke cases (incidence rate: 474/100,000 person-years). Ultimately, 13 predictor variables were included. RF achieved the highest accuracy (0.935), precision (0.923), sensitivity (recall: 0.947), and F1 score (0.935). Model evaluation demonstrated superior predictive performance of ML algorithms over conventional LR, with training/validation area under the curve (AUC)s of 0.777/0.779 (LR), 0.921/0.918 (BPNN), 0.988/0.980 (RF), 0.980/0.955 (DT), and 0.962/0.958 (XGBoost). Calibration analysis revealed a better fit for DT, LR and BPNN compared to RF and XGBoost model. Based on the optimal performance of the RF model, the ranking of factors in descending order of importance was: hypertension, age, diabetes, systolic blood pressure, waist, high-density lipoprotein Cholesterol, fasting blood glucose, physical activity, BMI, low-density lipoprotein cholesterol, total cholesterol, dietary habits, and family history of stroke. Using Youden's index as the optimal cutoff, the RF model stratified individuals into high-risk (>0.789) and low-risk (≤0.789) groups with robust discrimination.
Conclusions:
The ML-based prediction models demonstrated superior performance metrics compared to conventional LR and the RF is the optimal prediction model, providing an effective tool for risk stratification in primary stroke prevention in community settings.
Related Concept Videos
Steps in Outbreak Investigation
Statistical Software for Data Analysis and Clinical Trials

