Consistency of ranking was evaluated as new measure for prediction model stability: longitudinal cohort study

Yan Li1, Matthew Sperrin2, Darren M Ashcroft3

  • 1Health e-Research Centre, School of Health Sciences, Faculty of Biology, Medicine and Health, the University of Manchester, Manchester, Oxford Road, Manchester, M13 9PL, UK; School of Mathematical Sciences, Xiamen University, Xiamen, 361005, People's Republic of China.

Insights

Ranking consistency offers a novel way to evaluate individual patient risk prediction models, complementing traditional population-level assessments. This approach enhances the stability assessment of clinical prediction tools for cardiovascular disease.

Area of Science:

  • Clinical Epidemiology
  • Biostatistics
  • Machine Learning in Healthcare

Background:

  • Clinical risk prediction models are typically evaluated at the population level.
  • There is a need for measures assessing individual patient risk prediction stability.
  • Ranking stability offers a new metric for individual-level model evaluation.

Purpose of the Study:

  • To evaluate the utility of ranking as a measure for assessing individual-level stability of risk prediction models.
  • To compare ranking stability with absolute risk stability across different risk strata.

Main Methods:

  • Utilized a large patient cohort (3.66 million) from the Clinical Practice Research Datalink.
  • Examined 15 cardiovascular disease risk prediction models (machine learning and statistical).
  • Assessed consistency in ranking and absolute risk differences at individual patient level.

Main Results:

  • Population-level performance (C-statistics) was similar across models (~0.88).
  • At high absolute risk, models showed consistent ranking but inconsistent absolute risk.
  • At low absolute risk, models showed inconsistent ranking but more consistent absolute risk.

Conclusions:

  • Ranking consistency provides valuable, complementary information to absolute risk stability for individual risk prediction models.
  • Existing model development guidelines (e.g., TRIPOD, PROBAST) should integrate ranking assessment for individual-level stability.
  • This enhances the robustness and reliability of clinical risk prediction tools.
Abstract

Related Concept Videos

Longitudinal Research02:20

Longitudinal Research

Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
12.8K
Longitudinal Studies01:26

Longitudinal Studies

Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
295
Regression Toward the Mean01:52

Regression Toward the Mean

Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.6K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.2K
Ranks01:02

Ranks

Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
317
The Mantel-Cox Log-Rank Test01:19

The Mantel-Cox Log-Rank Test

The Mantel-Cox log-rank test is a widely used statistical method for comparing the survival distributions of two groups. It tests whether a statistically significant difference exists in survival times between the groups without assuming a specific distribution for the survival data, making it a non-parametric test. This flexibility makes the log-rank test particularly valuable in medical research and other fields where the timing of an event, such as death or disease recurrence, is of...
703