Related Experiment Video
Updated: Sep 3, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Prediction performance and fairness heterogeneity in cardiovascular risk models.
Uri Kartoun1, Shaan Khurshid2,3, Bum Chul Kwon1
1Center for Computational Health, IBM Research, 314 Main St., Cambridge, MA, 02142, USA.
Cardiovascular disease risk models show varied accuracy across patient subgroups. Performance declines with age and differs between sexes, necessitating subgroup analysis for equitable risk assessment.
Area of Science:
- Cardiovascular epidemiology
- Health informatics
- Biostatistics
Background:
- Clinical prediction models are crucial for cardiovascular disease (CVD) risk assessment, diagnosis, and management.
- Existing models may exhibit performance disparities across diverse demographic and clinical subgroups.
- Understanding this heterogeneity is vital for ensuring equitable healthcare outcomes.
Purpose of the Study:
- To investigate the heterogeneity in accuracy and fairness metrics of CVD risk prediction models.
- To evaluate performance variations across subpopulations defined by age, sex, and comorbidities.
- To assess two common CVD risk scores: the Cohorts for Heart and Aging in Genomic Epidemiology Atrial Fibrillation (CHARGE-AF) score and the Pooled Cohort Equations (PCE).
Main Methods:
- Calculated CHARGE-AF and PCE scores in three large datasets: Explorys, Mass General Brigham (MGB), and UK Biobank.
- Analyzed model performance using discrimination (concordance index) and calibration metrics.
- Examined subgroup performance based on age, sex, and presence of pre-existing conditions.
Main Results:
- Significant performance heterogeneity was observed across age, sex, and comorbidity subgroups for both CHARGE-AF and PCE scores.
- Model discrimination decreased with increasing age, with notable differences (e.g., CHARGE-AF concordance index from 0.72 to 0.57 in Explorys).
- Fairness metrics revealed considerable disparities, such as statistical parity differences between males and females, even when sex was not a predictor variable. Weak discrimination (<0.7) and suboptimal calibration were found in substantial population subsets (e.g., >75 years).
Conclusions:
- Clinical risk models exhibit substantial performance variations across patient subpopulations.
- Age, sex, and comorbidities significantly impact model accuracy and fairness.
- There is a critical need to characterize and quantify model behavior within specific subpopulations to ensure accurate, consistent, and equitable CVD risk assessment.
Related Concept Videos
Bias in Epidemiological Studies
Relative Risk
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Mechanistic Models: Compartment Models in Individual and Population Analysis
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.

