Related Experiment Video
Updated: Apr 29, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Discriminating clinical prediction models: how outcome incidence and decision thresholds shape clinical utility and
Farid Foroutan1, Jessica Weiss1, Martin Mayer2
1Ted Rogers Centre for Heart Research, University Health Network, Toronto, Ontario, Canada.
Objectives:
Clinical prediction models that estimate an individual's outcome risk are often evaluated using discrimination metrics, particularly the area under the receiver-operating characteristic curve (AUROC; c-statistic). Yet what constitutes a "good" AUROC is unclear because discrimination is often interpreted without explicit reference to outcome incidence or the risk thresholds that guide clinical decisions. We aimed to show that the clinical utility associated with a given AUROC depends jointly on decision thresholds and the outcome incidence in the target population.
Methods:
We conducted a simulation study and applied the findings to real examples. In simulations, we generated binary outcomes for 10,000 patients across 100 scenarios defined by combinations of target AUROC values (0.50-0.95) and overall outcome incidence (0.05-0.50), assuming perfect calibration. For each scenario, we evaluated clinical utility as net benefit across six treatment decision thresholds (5%-30%) and examined the resulting distribution of predicted risks and the proportion of patients exceeding each threshold. We then applied this framework to outcomes from a living clinical practice guideline on treatments for type 2 diabetes.
Results:
The simulations and applied examples showed that, for a well-calibrated model, what constitutes a "good" AUROC depends jointly on outcome incidence and the decision thresholds used in practice. As AUROC approached 0.5, predicted risks clustered increasingly tightly around the population incidence, limiting threshold-based decision-making unless the threshold was close to the incidence (so that some patients still fell above and some below it). Higher AUROC values widened the distribution of predicted risks, but clinical utility remained contingent on the proximity of thresholds to the incidence: thresholds far from the incidence required substantially higher discrimination for patients to cross the threshold and yield benefit.
Conclusion:
Whether a clinical prediction model's AUROC is "good" cannot be judged in isolation. Clinical usefulness arises from the joint relationship between discrimination, outcome incidence, and decision thresholds: For a well-calibrated model, AUROC determines the width of predicted risks, incidence anchors their center, and thresholds determine whether that width translates into actionable decisions. Under poor calibration, predicted risks may be dispersed in ways that do not reflect true risk separation, and these relationships may not hold. Accordingly, even well-calibrated models with high AUROC may have limited utility when thresholds are far from the outcome incidence, whereas models with modest AUROC may still be useful when thresholds lie near the incidence. Researchers should account for incidence and decision thresholds when designing prediction model studies and when interpreting the practical usefulness of a given AUROC value.
Related Concept Videos
Receiver Operating Characteristic Plot
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Pharmacodynamic Models: Overview
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Cancer Survival Analysis
