Related Experiment Video
Updated: Mar 15, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Community medical centers struggle to produce well-calibrated clinical prediction models: Data augmentation can help
Katherine E Brown1, Bradley A Malin1, Sharon E Davis1
1Vanderbilt University Medical Center Department of Biomedical Informatics, 2525 West End Avenue Suite 1400, Nashville, 37203, TN, USA.
Objective:
Machine learning models (ML) often require localization to perform optimally in local populations. We hypothesize that smaller community healthcare centers may not have the necessary patient volume to facilitate localization based on statistical guidelines. This work investigates the ability for community medical centers to localize ML and performs a simulation study to evaluate synthetic data generation (SDG) to augment local data for recalibration.
Methods:
We conducted an experiment using data from a real network of hospitals (two rural, one urban academic medical center) to predict 30-day unplanned hospital readmission and using data from a multi-site ICU dataset to simulate using synthetic data generation (SDG) in a network of hospitals of various sizes. We also performed a simulation study using data from a multi-site ICU dataset to evaluate the utility of SDG to augment local data volumes.
Results:
In the real-world evaluation, the urban medical center met the guidelines for the number of samples for recalibration (Required: 14,224, Available: 42,303) and had the best calibrated model using local data (α=0.1,β=1.05; best: α=0,β=1). For the smaller sites, neither site had the samples required for recalibration (Site 1: Required: 16461, Available: 3187; Site 2: Required: 15299, Available: 905). In the simulation study, deep learning-based SDG was most effective at improving calibration performance.
Conclusions:
Connections to large medical centers are not enough to promote accurate ML at all sites within a healthcare system. Data augmentation and SDG may provide the necessary data volumes to enable local recalibration at smaller facilities.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Clearance Models: Noncompartmental Models
The noncompartmental approach capitalizes on extensive sampling data, correlating the volume of distribution to systemic exposure and the administered dosage. This method enables...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...

