Related Experiment Videos
Improving 10-year cardiovascular disease risk prediction using automated machine learning
Simin He1, Juping Wang2, Le Zhao3
1Department of Health Statistics and Epidemiology, School of Public Health, MOE Key Laboratory of Coal Environmental Pathogenicity and Prevention, Shanxi Medical University, Taiyuan 030001, China; Department of Medical Information Retrieval, School of Management, Shanxi Medical University, Taiyuan 030001, China; Shanxi Provincial Key Laboratory of Major Diseases Risk Assessment, Shanxi Medical University, Taiyuan 030001, China.
Aims:
To develop a cardiovascular disease (CVD) risk prediction model with improved accuracy and interpretability by integrating diverse risk factors and applying Automated Machine Learning (AutoML), thereby enhancing clinical utility over conventional models.
Methods:
This is a prospective cohort study. Data were obtained from the Multi-Ethnic Study of Atherosclerosis (MESA), including baseline and fifth follow-up visits, comprising 4713 participants. Exercise and dietary data were harmonized via Metabolic Equivalent of Task (MET) and Healthy Eating Index-2015 (HEI-2015), respectively. Predictor selection was performed using the Boruta algorithm alongside Random Forest (RF) error rate cross-validation. Logistic regression, four traditional machine learning algorithms, and H2O AutoML were each applied for model training and evaluation. Finally, the best-performing model was further interpreted using SHapley Additive exPlanations (SHAP).
Results:
A total of 21 predictors were selected, including age, sex, and Total Cholesterol (TC). Among the evaluated models, H2O AutoML outperformed other methods with an accuracy of 0.864, specificity of 0.892, precision of 0.610, F1 score of 0.670, and a Youden index of 0.635, achieving the highest AUC of 0.882 (0.846-0.918). SHAP analysis revealed the relative importance of predictors, with age, TC and Digit Symbol Score (DSS) ranking highest.
Conclusions:
This study developed an AutoML-based CVD risk prediction model with superior discrimination and calibration, providing clinicians a practical tool for risk stratification. By enabling personalized prevention and early identification of high-risk individuals, this model has the potential to reduce CVD burden at the population level. Notably, DSS exhibited high importance and may represent a candidate risk marker.