Related Experiment Video
Updated: Jun 24, 2026

Determining the Functional Status of the Corticospinal Tract Within One Week of Stroke
Published on: February 22, 2020
Machine learning based prediction models for first stroke in community primary care: a systematic review and
Zijiao Zhang1, Yuru Zhang1, Zheng Li2
1School of Public Health, Southwest Medical University, No.1, Section 1, Xianglin Road, Longmatan District, Sichuan, 646000, China.
Insights
Machine learning models for first-stroke risk prediction show no significant advantage over traditional methods in community healthcare. Further research is needed for standardized validation and real-world clinical analysis to improve stroke prevention.
Area of Science:
- Medical Informatics
- Public Health
- Cardiovascular Research
Background:
- Stroke is a leading global cause of death and disability, with a rising burden in low- and middle-income countries.
- Accurate prediction of first-stroke events is critical for effective prevention strategies.
- The clinical utility of artificial intelligence (AI) and machine learning (ML) models in primary community healthcare for stroke risk prediction remains largely unverified.
Purpose of the Study:
- To systematically review and analyze the application of algorithms, particularly ML-based ones, in first-stroke risk prediction models within community settings.
- To assess the performance and clinical utility of various predictive models in primary care.
Main Methods:
- A comprehensive literature search was conducted in PubMed, EMBASE, and the Cochrane Library up to June 30, 2024.
- Studies developing or validating multivariable stroke risk models for primary care were included.
- Meta-analysis of AUC/C-statistic values with 95% confidence intervals was performed to evaluate model performance.
Main Results:
- A total of 43 studies with 93 models were analyzed; Cox regression was the most common traditional method.
- Machine learning models, including logistic regression, random forest, and eXtreme Gradient Boosting, showed comparable performance to traditional models (pooled AUCs ranging from 0.76 to 0.77).
- A high risk of bias (76.7%) and concerns regarding applicability to community settings (20.9%) were noted in the included studies; ML models did not outperform traditional regression models.
Conclusions:
- The clinical utility of ML models for first-stroke risk prediction in community primary healthcare is currently unverified.
- Current ML models offer no significant advantage over traditional regression methods and face challenges in innovation, standardization, and validation.
- Methodological flaws and applicability concerns necessitate caution; future research should prioritize standardized validation and real-world clinical analysis for effective stroke prevention tools.
Background:
Stroke is the third leading cause of death and the fourth leading cause of disability globally, particularly in low- and middle-income countries, where the disease burden is increasing significantly. Risk prediction of first-stroke is crucial for stroke prevention. Although artificial intelligence (AI)-based prediction models have rapidly developed in recent years, their clinical utility in primary community healthcare settings remains unclear. This study aimed to systematically review and analyze the application of algorithms, especially machine learning (ML)-based ones, in first-stroke risk prediction models within community settings.
Methods:
We searched PubMed, EMBASE, and the Cochrane Library from inception to 30 June 2024 for studies developing or validating multivariable stroke risk models for primary care. We analyzed AUC/C-statistic values with 95% confidence intervals (CI) and conducted a meta-analysis.
Results:
A total of 43 studies encompassing 93 models were included. The Cox regression model was the most common traditional method (n = 32, 74.42% of traditional models). Among ML models, 25 algorithms were identified, with logistic regression being the most frequent (n = 8, 18.60%). The meta-analysis showed a pooled AUC of 0.79 (95% CI: 0.77-0.81) for the Cox regression model, while logistic regression, random forest, and eXtreme Gradient Boosting models yielded 0.76 (95% CI: 0.64-0.89), 0.77 (95% CI: 0.61-0.93), and 0.77 (95% CI: 0.59-0.94), respectively. Notably, 76.7% of the included studies had a high risk of bias, and 20.9% raised high concerns regarding their applicability to community primary healthcare settings. ML models did not outperform traditional regression models, with significant performance variability observed.
Conclusion:
Despite the increasing use of ML models for first-stroke risk prediction, their clinical utility in community primary healthcare remains unverified. Current ML models show no significant advantage over traditional regression methods and face challenges in algorithm innovation, data standardization, and external validation. Furthermore, the prevalence of methodological flaws and applicability concerns among existing studies underscores the need for caution. Future research should focus on standardized validation and real-world clinical analysis to develop effective stroke prevention tools.