Related Experiment Video
Updated: Nov 20, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Developing a Clinical Prediction Score: Comparing Prediction Accuracy of Integer Scores to Statistical Regression
Vigneshwar Subramanian1, Edward J Mascha2, Michael W Kattan3
1From the Cleveland Clinic Lerner College of Medicine at Case Western Reserve University, Cleveland, Ohio.
Simplifying regression models into integer scores significantly reduces prediction accuracy (AUC and IPA). Retaining continuous variables in regression models is crucial for precise clinical predictions. Always validate the intended clinical model.
Area of Science:
- Biostatistics
- Clinical Prediction Modeling
- Health Informatics
Background:
- Statistical regression models are often simplified into integer scores or risk classifications for ease of use.
- This simplification process can lead to a loss of valuable information and reduced prediction accuracy.
- The impact of such simplification on model performance requires thorough investigation.
Purpose of the Study:
- To investigate the impact of simplifying regression models into integer scores on prediction accuracy.
- To compare the performance of continuous regression models versus simplified integer-score models.
- To assess the effects of dichotomization and risk stratification on model discrimination and calibration.
Main Methods:
- A simulation study with logistic regression models and continuous covariates was conducted.
- Simulated data sets (n=1000) were generated with target Area Under the Curve (AUC) values of 0.7, 0.8, or 0.9.
- Continuous variables were dichotomized, and models were refit to create integer scores and risk classifications. Performance was evaluated using AUC and Index of Prediction Accuracy (IPA).
- An external clinical data set was used to validate findings.
Main Results:
- Logistic regression models using continuous covariates consistently outperformed simplified integer-score models.
- Simplification to integer scores resulted in an average decrease of 0.057-0.094 in AUC and 6.2%-17.5% in IPA in simulations.
- The largest performance decrease occurred during the dichotomization step, which also increased model optimism.
- External validation on a clinical data set showed a decrease of 0.06 in AUC and 13% in IPA when converting to an integer score.
Conclusions:
- Converting regression models to integer scores and risk classification systems considerably decreases model performance.
- Retaining continuous variables in regression models is recommended for maximizing prediction accuracy.
- Researchers should use unaltered regression models for individual patient predictions and ensure proper validation of the intended clinical model.
More Related Videos
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Receiver Operating Characteristic Plot
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Regression Toward the Mean
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...

