Related Experiment Video
Updated: Jan 21, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Do population-level risk prediction models that use routinely collected health data reliably predict individual
Yan Li1, Matthew Sperrin1, Miguel Belmonte1
1Health e-Research Centre, School of Health Sciences, Faculty of Biology, Medicine and Health, The University of Manchester, Manchester Academic Health Sciences Centre (MAHSC), Oxford Road, Manchester, M13 9PL, UK.
Insights
Routinely collected health data show significant variation across clinical sites, impacting individual cardiovascular disease (CVD) risk predictions. While models perform well for populations, individual predictions carry substantial uncertainty that clinicians and patients must acknowledge.
Area of Science:
- Health Informatics
- Epidemiology
- Biostatistics
Background:
- Routinely collected health data (RCD) are increasingly used for clinical research and risk prediction.
- Heterogeneity in data recording and patient populations across clinical sites poses challenges for reliable risk prediction.
- Existing risk prediction models may not fully account for site-specific variations.
Purpose of the Study:
- To assess the reliability of individual risk predictions derived from RCD, considering inter-site heterogeneity.
- To evaluate the performance of a random effects model compared to standard QRISK3 predictions for cardiovascular disease (CVD).
Main Methods:
- Utilized a large dataset of 3.6 million patients from 392 sites in the Clinical Practice Research Datalink.
- Employed Cox models incorporating QRISK3 predictors and a frailty (random effect) term for each site to address unmeasured site variability.
- Analyzed variations in data recording (e.g., BMI missingness) and CVD incidence rates across general practices.
Main Results:
- Significant variation observed in data recording completeness (BMI missingness: 18.7%-60.1%) and CVD incidence rates (0.4-1.3 per 100 patient-years) between practices.
- Individual CVD risk predictions from the random effects model showed inconsistency with QRISK3 predictions (e.g., 10% QRISK3 risk had a 95% range of 7.2%-13.7% with random effects).
- Random variability accounted for only a small portion of the observed prediction inconsistency; however, the random effects model demonstrated equivalent discrimination and calibration to QRISK3.
Conclusions:
- Risk prediction models using RCD perform adequately at a population level but exhibit considerable uncertainty for individual predictions.
- The heterogeneity between clinical sites in data and populations significantly affects the reliability of individual risk estimates.
- Clinicians and patients require clear communication regarding the inherent uncertainty in individual risk predictions derived from routinely collected data.
Abstract:
The objective of this study was to assess the reliability of individual risk predictions based on routinely collected data considering the heterogeneity between clinical sites in data and populations. Cardiovascular disease (CVD) risk prediction with QRISK3 was used as exemplar. The study included 3.6 million patients in 392 sites from the Clinical Practice Research Datalink. Cox models with QRISK3 predictors and a frailty (random effect) term for each site were used to incorporate unmeasured site variability. There was considerable variation in data recording between general practices (missingness of body mass index ranged from 18.7% to 60.1%). Incidence rates varied considerably between practices (from 0.4 to 1.3 CVD events per 100 patient-years). Individual CVD risk predictions with the random effect model were inconsistent with the QRISK3 predictions. For patients with QRISK3 predicted risk of 10%, the 95% range of predicted risks were between 7.2% and 13.7% with the random effects model. Random variability only explained a small part of this. The random effects model was equivalent to QRISK3 for discrimination and calibration. Risk prediction models based on routinely collected health data perform well for populations but with great uncertainty for individuals. Clinicians and patients need to understand this uncertainty.
Related Concept Videos
Predicting Molecular Geometry
Relative Risk
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Factors Affecting the Risk of Infection
The integrity and count of the white blood cells help the body resist pathogens and fight infection. When impaired, it reduces the body's resistance to pathogens. The acidic pH levels of the gastrointestinal, genitourinary tracts, and skin...
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...

