Related Experiment Video
Updated: May 23, 2025

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Interpretable machine learning leverages proteomics to improve cardiovascular disease risk prediction and biomarker
Héctor Climente-González1, Min Oh2, Urszula Chajewska2
1Human Genetics Centre of Excellence, Novo Nordisk Research Centre Oxford, The Innovation Building, Roosevelt Dr, Headington, Oxford, OX3 7FZ, United Kingdom. HECG@novonordisk.com.
Insights
Predicting cardiovascular disease (CVD) risk is vital. Using UK Biobank proteomics data with machine learning, this study developed a more accurate CVD risk prediction model than traditional scores.
Area of Science:
- Genomics and Proteomics
- Biomedical Informatics
- Cardiovascular Medicine
Background:
- Cardiovascular diseases (CVDs) are a leading cause of mortality and disability.
- Accurate CVD risk prediction is essential for prevention and early intervention.
- UK Biobank Proteomics data offers a novel resource for disease association studies.
Purpose of the Study:
- To predict 10-year CVD risk using proteomics data and clinical risk factors.
- To develop an interpretable machine learning model for CVD risk assessment.
- To identify potential therapeutic targets through gene association.
Main Methods:
- Utilized UK Biobank Pharma Proteomics Project data from 50,057 participants (aged 40-69).
- Employed Explainable Boosting Machine (EBM), an interpretable ML model.
- Included 2923 proteins and 55 clinical risk factors as features; evaluated using 10-fold cross-validation.
Main Results:
- The EBM proteomics model achieved an AUROC of 0.767 and AUPRC of 0.241, outperforming existing risk scores.
- Incorporating clinical features improved performance to AUROC 0.785 and AUPRC 0.284.
- Demonstrated consistent model performance across diverse sexes and ethnicities.
Conclusions:
- Developed a more accurate and explanatory framework for proteomics data analysis in CVD risk prediction.
- The approach supports individualized disease risk prediction.
- Facilitates the identification of target genes for future drug development.
Background:
Cardiovascular diseases (CVDs) rank amongst the leading causes of long-term disability and mortality. Predicting CVD risk and identifying associated genes are crucial for prevention, early intervention, and drug discovery. The recent availability of UK Biobank Proteomics data enables investigation of blood proteins and their association with a variety of diseases. We sought to predict 10 year CVD risk using this data modality and known CVD risk factors.
Methods:
We focused on the UK Biobank participants that were included in the UK Biobank Pharma Proteomics Project. After applying exclusions, 50,057 participants were included, aged 40-69 years at recruitment. We employed the Explainable Boosting Machine (EBM), an interpretable machine learning model, to predict the 10 year risk of primary coronary artery disease, ischemic stroke or myocardial infarction. The model had access to 2978 features (2923 proteins and 55 risk factors). Model performance was evaluated using 10-fold cross-validation.
Results:
The EBM model using proteomics outperforms equation-based risk scores such as PREVENT, with a receiver operating characteristic curve (AUROC) of 0.767 and an area under the precision-recall curve (AUPRC) of 0.241; adding clinical features improves these figures to 0.785 and 0.284, respectively. Our models demonstrate consistent performance across sexes and ethnicities and provide insights into individualized disease risk predictions and underlying disease biology.
Conclusions:
In conclusion, we present a more accurate and explanatory framework for proteomics data analysis, supporting future approaches that prioritize individualized disease risk prediction, and identification of target genes for drug development.
Related Concept Videos
Proteomics
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Blood Studies for Cardiovascular System II: CRP, Hcy, and Cardiac Natriuretic Peptide Markers
These markers indicate stress or strain on the heart muscle:
Natriuretic Peptides (BNP)
Cardiac myocytes produce these hormones in response to ventricular stretching...
Blood Studies for Cardiovascular System I: Cardiac Biomarkers
The essential diagnostic tools for detecting myocardial necrosis and monitoring individuals suspected of having acute coronary syndrome (ACS) include:
Troponins
Troponins, particularly cardiac troponins I and T, are the most precise and sensitive markers of myocardial injury. They are detectable within 4-6 hours of myocardial injury and remain...

