Related Experiment Video
Updated: Jun 16, 2026

Predicting Amputation using Local Circulating Mononuclear Progenitor Cells in Angioplasty-treated Patients with Critical Limb Ischemia
Published on: September 22, 2020
Machine Learning-Driven Prediction of Coronary Artery Disease Risk Based on UK Biobank Plasma Proteomics
Yuezhong Huang1,2, Xiaoli Chen1,2, Hao Zhang1,2
1Zhejiang Provincial Clinical Research Center for Pediatric Precision Medicine The Second Affiliated Hospital and Yuying Children's Hospital of Wenzhou Medical University Wenzhou Zhejiang China.
Insights
Integrating polygenic risk scores and proteomic data with conventional factors significantly enhances coronary artery disease (CAD) risk prediction. This approach offers improved precision cardiovascular medicine and simplified risk stratification tools for better patient outcomes.
Area of Science:
- Cardiovascular Medicine
- Genomics
- Proteomics
- Biomarker Discovery
Background:
- Coronary artery disease (CAD) remains a leading cause of mortality globally.
- Conventional risk models for CAD exhibit limited predictive accuracy.
- There is a need for integrated approaches combining diverse data types for improved risk prediction.
Purpose of the Study:
- To develop and validate a unified model for enhanced coronary artery disease (CAD) risk prediction.
- To integrate conventional risk factors, polygenic risk scores, and large-scale proteomics data.
- To assess the incremental predictive value of proteomic data in CAD risk stratification.
Main Methods:
- Utilized UK Biobank data, including plasma proteomics and genetic risk data, excluding prevalent CAD cases.
- Trained CatBoost models incorporating conventional risk factors, polygenic risk scores, and a 202-protein proteomic risk score.
- Employed least absolute shrinkage and selection operator (LASSO) Cox regression for risk score derivation and Shapley Additive Explanations (SHAP) for feature selection, identifying a 9-protein panel.
Main Results:
- The proteomic risk score demonstrated a dose-dependent association with CAD risk across validation cohorts.
- Integration of polygenic and proteomic risk scores significantly improved CAD risk discrimination compared to conventional factors alone (AUC increased from 0.750 to 0.789 in internal validation).
- A compact 9-protein panel, including GDF15, MMP12, and ACE2, captured substantial proteomic predictive information.
Conclusions:
- Integrating conventional risk factors, polygenic risk scores, and proteomic data substantially enhances CAD risk prediction.
- Proteomics plays a crucial role in advancing precision cardiovascular medicine.
- The developed models and identified protein panels offer potential for simplified and more accurate cardiovascular risk stratification tools.
Background:
Coronary artery disease (CAD) is a leading global cause of mortality, yet the predictive accuracy of conventional risk models is limited. Here, we integrate conventional risk factors, polygenic risk scores, and large-scale proteomics to develop a unified model for enhanced CAD risk prediction.
Methods:
Using data from UK Biobank, participants with plasma proteomics and genetic risk data were included after excluding prevalent CAD. Participants from England were split into training (n=32 330) and internal validation (n=13 857) sets, and Scotland/Wales participants formed an external validation set (n=5775). Incident CAD was ascertained from linked health records. A 202-protein proteomic risk score was derived by least absolute shrinkage and selection operator Cox regression, and CatBoost models were trained using conventional risk factors alone and with incremental addition of polygenic risk scores and protein proteomic risk scores; Shapley Additive Explanations-guided forward selection identified a compact protein panel.
Results:
Across cohorts, the median age was 58 years and ∼45% were men. Protein proteomic risk score was dose-dependently associated with CAD risk. Compared with conventional risk factors alone, integrating polygenic risk scores and protein proteomic risk scores improved discrimination, with the area under the curve increasing from 0.750 (95% CI, 0.732-0.767) to 0.789 (95% CI, 0.772-0.805) in internal validation and from 0.717 (95% CI, 0.683-0.750) to 0.762 (95% CI, 0.732-0.791) in external validation. A 9-protein panel (GDF15 [growth differentiation factor 15], MMP12 [matrix metalloproteinase 12], NPPB [natriuretic peptide B], PGF [placental growth factor], REN [renin], ADGRG2 [adhesion G-protein coupled receptor], ACE2 [angiotensin-converting enzyme 2], CDCP1 [CUB domain-containing protein 1], CXCL17 [C-X-C motif chemokine ligand 17)]) captured most proteomic predictive information.
Conclusions:
Our findings demonstrate that integrating conventional risk factors, polygenic risk scores, and proteomic data improves CAD risk prediction. This study highlights the utility of proteomics in precision cardiovascular medicine and simplified risk stratification tools.
Related Concept Videos
Blood Studies for Cardiovascular System I: Cardiac Biomarkers
The essential diagnostic tools for detecting myocardial necrosis and monitoring individuals suspected of having acute coronary syndrome (ACS) include:
Troponins
Troponins, particularly cardiac troponins I and T, are the most precise and sensitive markers of myocardial injury. They are detectable within 4-6 hours of myocardial injury and remain...
Pharmacogenomics: Identification of New Drug Targets
Blood Studies for Cardiovascular System II: CRP, Hcy, and Cardiac Natriuretic Peptide Markers
These markers indicate stress or strain on the heart muscle:
Natriuretic Peptides (BNP)
Cardiac myocytes produce these hormones in response to ventricular stretching...
Coronary Artery Disease I: Introduction