Related Experiment Video
Updated: Jul 19, 2026

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
Machine learning to predict stroke risk from routine hospital data: A systematic review
William Heseltine-Carp1, Megan Courtman2, Daniel Browning1
1University of Plymouth, Room N6, ITTC Building, Plymouth Science Park, Plymouth PL68BX, UK.
Insights
Machine learning (ML) shows promise for predicting stroke risk using hospital data, outperforming traditional methods. However, further research is needed to improve model accuracy and clinical integration.
Area of Science:
- Cardiology
- Medical Informatics
- Data Science
Background:
- Stroke is a major cause of death and disability.
- Current stroke risk prediction tools have limited accuracy, especially for non-atrial fibrillation patients.
- There is a need for improved stroke risk stratification models.
Purpose of the Study:
- To systematically review research on machine learning (ML) for stroke risk prediction using routine hospital data.
- To identify methodological limitations and provide recommendations for future ML-based stroke prediction research.
Main Methods:
- A systematic review of 49 original research articles published between January 2013 and December 2024.
- Studies utilized machine learning algorithms and routine hospital data to predict stroke risk.
- Searches were conducted in the PUBMED database, including general and atrial fibrillation-specific populations.
Main Results:
- Machine learning models demonstrated high accuracy (AUCs 0.64-0.99) in predicting stroke risk.
- ML models consistently outperformed traditional tools like CHA 2 DS 2 -VASc.
- ML identified novel risk factors from ECG, lab, and echocardiography data, but dataset quality, overfitting, and lack of validation were noted.
Conclusions:
- Machine learning holds significant potential for stroke risk prediction and novel risk factor identification.
- Improvements in study methodology, including adherence to EQUATOR guidelines and interdisciplinary collaboration, are crucial.
- Prospective studies are needed to validate ML models and assess barriers to clinical integration.
Purpose:
Stroke remains a leading cause of morbidity and mortality. Despite this, current risk stratification tools such as CHA2DS2-VASc and QRISK3 are of limited accuracy, particularly in those without a diagnosis of atrial-fibrillation. Hence, there is a need for more accurate stroke risk prediction models. Machine-learning (ML) may provide a solution to this by leveraging existing routine hospital databases to build accurate stroke risk prediction models and identify novel risk factors for stroke.
Aims:
In this systematic review we appraise current research using ML to predict stroke risk from routine hospital data. Based on these findings we then highlight common methodological limitations and recommendations for future research.
Methods:
In this review we identify 49 original research (38 in the general population and 11 in AF specific populations) articles from the PUBMED database from January-2013 to December-2024 using ML and routine hospital data to predict the risk of stroke.
Results:
ML models were able to accurately predict stroke risk in both AF specific and general populations, with AUCs ranging from 0.64 to 0.99. Where tested, ML also consistently outperformed traditional risk stratification tool, such as CHA2DS2-VASc. ML also appeared useful in identifying several novel risk factors from electrocardiogram, laboratory test and echocardiography data. However, the quality of datasets were often limited, there was a high suspicion of overfitting and models often lacked calibration, external validation and explainability analysis.
Conclusion:
Whilst ML has shown great potential in stroke prediction and identifying novel risk factors for stroke, improvements in study methodology is required prior to integration of ML into routine healthcare. Future research should adhere to the EQUATOR guidance on prediction models and encourage interdisciplinary collaboration between computer scientists and clinicians. Further prospective RCTs are also required to validate models in the clinical setting and the identify barriers of integrating ML into routine healthcare.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020