Related Experiment Video
Updated: Aug 2, 2025

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Use of machine learning to identify risk factors for coronary artery disease
Alexander A Huang1,2, Samuel Y Huang1,3
1Department of Statistics and Data Science, Cornell University, Ithaca, New York, United States of America.
Abstract:
Coronary artery disease (CAD) is the leading cause of death in both developed and developing nations. The objective of this study was to identify risk factors for coronary artery disease through machine-learning and assess this methodology. A retrospective, cross-sectional cohort study using the publicly available National Health and Nutrition Examination Survey (NHANES) was conducted in patients who completed the demographic, dietary, exercise, and mental health questionnaire and had laboratory and physical exam data. Univariate logistic models, with CAD as the outcome, were used to identify covariates that were associated with CAD. Covariates that had a p<0.0001 on univariate analysis were included within the final machine-learning model. The machine learning model XGBoost was used due to its prevalence within the literature as well as its increased predictive accuracy in healthcare prediction. Model covariates were ranked according to the Cover statistic to identify risk factors for CAD. Shapely Additive Explanations (SHAP) explanations were utilized to visualize the relationship between these potential risk factors and CAD. Of the 7,929 patients that met the inclusion criteria in this study, 4,055 (51%) were female, 2,874 (49%) were male. The mean age was 49.2 (SD = 18.4), with 2,885 (36%) White patients, 2,144 (27%) Black patients, 1,639 (21%) Hispanic patients, and 1,261 (16%) patients of other race. A total of 338 (4.5%) of patients had coronary artery disease. These were fitted into the XGBoost model and an AUROC = 0.89, Sensitivity = 0.85, Specificity = 0.87 were observed (Fig 1). The top four highest ranked features by cover, a measure of the percentage contribution of the covariate to the overall model prediction, were age (Cover = 21.1%), Platelet count (Cover = 5.1%), family history of heart disease (Cover = 4.8%), and Total Cholesterol (Cover = 4.1%). Machine learning models can effectively predict coronary artery disease using demographic, laboratory, physical exam, and lifestyle covariates and identify key risk factors.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
06:16Signal Acquisition, Score Interpretation, and Economics of a Non-Invasive Point-of-Care Test for Coronary Artery Disease
Published on: August 9, 2024
Related Concept Videos
Coronary Artery Disease I: Introduction
Coronary Artery Disease IV: Preventive Measures
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT
Coronary Artery Disease II: Pathophysiology
Atherosclerosis II: Clinical Manifestations and Diagnostic Tests
Coronary Artery Disease V: Interprofessional Care