Related Experiment Video
Updated: Sep 13, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Correcting for case-mix shift when developing clinical prediction models
Haya Elayan1, Matthew Sperrin2, Glen P Martin2
1Division of Informatics, Imaging and Data Science, Faculty of Biology, Medicine and Health, University of Manchester, Manchester, UK. haya.elayan@postgrad.manchester.ac.uk.
A new Membership-based method effectively corrects for case-mix shift in clinical prediction models (CPMs), especially with limited target data. This approach improves model performance by re-weighting data to match the target population distribution.
Area of Science:
- Clinical Epidemiology
- Health Informatics
- Biostatistics
Background:
- Clinical prediction models (CPMs) can be affected by case-mix shift, where predictor distributions change in development datasets.
- This shift can impact model performance during deployment.
- This study leverages observed case-mix shifts within development data to address deployment-phase shifts.
Purpose of the Study:
- To introduce and evaluate a novel Membership-based method for correcting case-mix shift during CPM development.
- To assess the impact of this method on CPM predictive performance under various shift and sample size scenarios.
Main Methods:
- A Membership-based method using a probabilistic similarity metric to re-weight source data samples.
- Application to a real-world dataset of myocardial infarction patients with out-of-hospital cardiac arrest.
- Evaluation across nine scenarios, comparing the proposed method against models ignoring shift or using only recent data.
Main Results:
- The Membership-based method shows promise, particularly with insufficient target set sample sizes, achieving an optimism-adjusted calibration slope (c-slope) of 0.98 in partial shift scenarios.
- When target sample size was sufficient, the Unweighted model on target data only performed better (c-slope 0.95) than the Membership-based model (c-slope 0.92).
- In complete case-mix shift scenarios, both Membership-based and Unweighted models performed similarly, achieving c-slopes of 0.77 (insufficient target data) and 0.94 (sufficient target data).
Conclusions:
- The Membership-based method is a promising approach for addressing case-mix shift in CPM development, especially when target data is limited.
- Further research is needed to validate the method and explore its effectiveness with other types of data distribution shifts.
- Optimizing CPMs for evolving data distributions is crucial for maintaining predictive accuracy.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Bias in Epidemiological Studies
Confounding in Epidemiological Studies
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...