Related Experiment Video
Updated: Oct 30, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Consistency of ranking was evaluated as new measure for prediction model stability: longitudinal cohort study
Yan Li1, Matthew Sperrin2, Darren M Ashcroft3
1Health e-Research Centre, School of Health Sciences, Faculty of Biology, Medicine and Health, the University of Manchester, Manchester, Oxford Road, Manchester, M13 9PL, UK; School of Mathematical Sciences, Xiamen University, Xiamen, 361005, People's Republic of China.
Insights
Ranking consistency offers a novel way to evaluate individual patient risk prediction models, complementing traditional population-level assessments. This approach enhances the stability assessment of clinical prediction tools for cardiovascular disease.
Area of Science:
- Clinical Epidemiology
- Biostatistics
- Machine Learning in Healthcare
Background:
- Clinical risk prediction models are typically evaluated at the population level.
- There is a need for measures assessing individual patient risk prediction stability.
- Ranking stability offers a new metric for individual-level model evaluation.
Purpose of the Study:
- To evaluate the utility of ranking as a measure for assessing individual-level stability of risk prediction models.
- To compare ranking stability with absolute risk stability across different risk strata.
Main Methods:
- Utilized a large patient cohort (3.66 million) from the Clinical Practice Research Datalink.
- Examined 15 cardiovascular disease risk prediction models (machine learning and statistical).
- Assessed consistency in ranking and absolute risk differences at individual patient level.
Main Results:
- Population-level performance (C-statistics) was similar across models (~0.88).
- At high absolute risk, models showed consistent ranking but inconsistent absolute risk.
- At low absolute risk, models showed inconsistent ranking but more consistent absolute risk.
Conclusions:
- Ranking consistency provides valuable, complementary information to absolute risk stability for individual risk prediction models.
- Existing model development guidelines (e.g., TRIPOD, PROBAST) should integrate ranking assessment for individual-level stability.
- This enhances the robustness and reliability of clinical risk prediction tools.
Objective:
Clinical risk prediction models are generally assessed on population level with a lack of measures that evaluate their stability at predicting risks of individual patients. This study evaluated the use of ranking as a measure to assess individual level stability between risk prediction models.
Study Design And Setting:
A large patient cohort (3.66 million patients with 0.11 million cardiovascular events) extracted from the Clinical Practice Research Datalink was used in the exemplar of cardiovascular disease risk prediction.
Results:
It was found that 15 models (including machine learning and statistical models) had similar population-level model performance (C statistics about 0.88). For patients with high absolute risks, the models were more consistent in ranking of risk predictions (interquartile range (IQR) of differences in rank percentiles -0.6 to 1.0), but inconsistent in absolute risk (IQR of differences in absolute risk -18.8 to 9.0). At low risk, the reverse was true with inconsistent ranking but more consistent absolute risk.
Conclusion:
Consistency of ranking of individual risk predictions is a useful measure to assess risk prediction models providing complementary information to absolute risk stability. Model developing guidelines including "TRIPOD" and "PROBAST" should incorporate ranking to assess individual level stability between risk prediction models.
More Related Videos
Related Concept Videos
Longitudinal Research
Longitudinal Studies
Regression Toward the Mean
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Ranks
The Mantel-Cox Log-Rank Test

