Related Experiment Video
Updated: May 25, 2026

The WATCHMAN Left Atrial Appendage Closure Device for Atrial Fibrillation
Published on: February 28, 2012
Advancing stroke prevention in atrial fibrillation: a systematic review of machine learning-based risk prediction
Md Mohaimenul Islam1, Arinze Nkemdirim Okere1
1Division of Outcomes and Practice Advancement, Department of Pharmacy Practice, School of Pharmacy and Pharmaceutical Sciences, University at Buffalo, NY, USA; Institute for Artificial Intelligence and Data Science, University at Buffalo, NY, USA.
Insights
Machine learning models show potential for predicting ischemic stroke in atrial fibrillation (AF) patients, outperforming traditional scores. However, methodological limitations necessitate caution before clinical adoption.
Area of Science:
- Cardiology
- Medical Informatics
- Artificial Intelligence
Background:
- Atrial fibrillation (AF) significantly increases ischemic stroke risk, with current risk scores having modest predictive power.
- Electronic Health Record (EHR) data offers complex, high-dimensional information for improved risk stratification.
- Existing tools like CHA₂DS₂-VASc have limitations in capturing nonlinear interactions within EHR data.
Purpose of the Study:
- To systematically evaluate machine learning (ML) models for ischemic stroke prediction in AF patients using EHR data.
- To assess the predictive performance, methodological rigor, and clinical readiness of these ML models.
- To compare ML model performance against the established CHA₂DS₂-VASc score.
Main Methods:
- Systematic literature search across major databases (PubMed, Embase, Scopus, Web of Science) following PRISMA 2020 guidelines.
- Inclusion of studies developing or validating ML models for ischemic stroke prediction in AF patients using EHR data.
- Methodological quality assessment using PROBAST and TRIPOD-AI frameworks.
Main Results:
- Eight studies (2017-2024) including over 800,000 patients were analyzed.
- Supervised ensemble ML models generally outperformed the CHA₂DS₂-VASc score (AUROCs 0.66-0.91 vs. 0.54-0.68).
- Significant heterogeneity in performance, limited external validation, infrequent use of explainable AI, and high risk of bias (88% in analysis domain) were noted.
Conclusions:
- Claims of ML superiority over CHA₂DS₂-VASc require caution due to pervasive methodological limitations.
- Current evidence is insufficient for widespread clinical adoption of ML models for stroke risk prediction in AF.
- Future research needs rigorous external validation, longitudinal modeling, and prospective evaluation for clinical translation.
Background:
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and confers a four to fivefold increase in ischemic stroke risk, accounting for approximately 15 - 20% of all stroke events globally. Despite this burden, the predominant risk stratification tool, the CHA2DS2-VASc score, achieves only modest discrimination, constrained by its static, additive architecture that cannot capture the nonlinear, high-dimensional interactions inherent in real-world electronic health record (EHR) data. This evidence gap creates a dual clinical hazard: under-anticoagulation in high-risk patients and unnecessary bleeding exposure in those whose risk is overestimated. This study aimed to systematically evaluate the predictive performance, methodological rigor, and clinical readiness of machine learning (ML) models derived from EHR data for the prediction of ischemic stroke in patients with AF.
Methods:
A systematic search of PubMed, Embase, Scopus, and Web of Science was conducted from inception through September 2025, following PRISMA 2020 guidelines. Studies were eligible if they developed or validated ML models for ischemic stroke prediction using EHR data in adults with AF and reported at least one quantitative performance metric. Methodological quality was assessed using the PROBAST and TRIPOD-AI frameworks.
Results:
Eight studies (2017 to 2024) encompassing 809,523 patients across seven countries were included. Supervised ensemble methods consistently outperformed CHA2DS2-VASc, with AUROCs ranging from 0.66 to 0.91 versus 0.54 to 0.68 for the traditional score. However, performance varied substantially: several models achieved only marginal gains (AUROC 0.63 - 0.69), and the AUROC range reflects pronounced heterogeneity rather than uniform superiority. Critical barriers persist - only one study performed external validation; fewer than half applied explainable AI techniques; class imbalance was rarely addressed; and 88% of studies received a high risk of bias rating in the analysis domain under PROBAST, a finding that substantially limits confidence in the reported performance estimates.
Conclusion:
In light of the pervasive methodological limitations identified, including high analytic risk of bias, absence of external validation, and lack of model interpretability, claims of ML superiority over CHA2DS2-VASc must be interpreted with caution. While ML models demonstrate potential discriminative improvements, current evidence is insufficient to support clinical adoption. Translating algorithmic promise into bedside impact requires dynamic longitudinal modeling, rigorous multisite external validation, transparent risk attribution, and prospective evaluation within real-world EHR workflows.