Related Experiment Video
Updated: Apr 4, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Validation of a Risk-Prediction Model in the Presence of Outcome Misclassification
Runjia Zou1, Brian D Williamson1,2,3, Susan M Shortreed1,2
1Department of Biostatistics, University of Washington, Seattle, Washington, USA.
Abstract:
Electronic health records (EHRs) provide a rich data source for building prediction models to improve the quality of care. However, EHR data are prone to measurement error, and outcomes used to evaluate prediction model performance may be misclassified. Comparing risk predictions to misclassified outcomes will result in unreliable estimates of a prediction model's performance. We propose a method that leverages a smaller chart review sample with gold-standard outcome measurements to adjust validation for outcome misclassification and provide more accurate prediction model evaluation. We derive formulae to estimate true positive rate (TPR), false positive rate (FPR), positive predictive value (PPV), negative predictive value (NPV), and area under the receiver operating characteristic curve (AUC) in the presence of outcome misclassification. Different scenarios of misclassification are explored, including when the misclassification is independent or dependent on features, and when misclassification is unidirectional (e.g., only missed diagnoses) or bidirectional. In simulation studies, we compare the bias and 95% confidence interval coverage of performance estimates obtained using our proposed method to those estimated with misclassified outcomes (without accounting for misclassification) or in the smaller chart review sample. Simulation results indicate that, across all misclassification scenarios examined, our proposed estimates have good accuracy and improved precision. Outcome misclassification should be considered when evaluating a prediction model's performance in order to accurately inform decision-making about whether and how to use a clinical prediction model.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Receiver Operating Characteristic Plot
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Mechanistic Models: Compartment Models in Individual and Population Analysis

