Related Experiment Video
Updated: Sep 11, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Testing Transferability of a Mortality Risk Model for Interhospital Transfer Using Real-World Electronic Health
Rachel A Hadler1,2,3,4, Sahithi Krishnaveni Lakamana5, Tristan Moorman6
1Department of Anesthesiology, Division of Critical Care, Emory University School of Medicine, Atlanta, GA.
Background:
Mortality risk is often uncertain at the time of interhospital transfer (IHT).
Objectives:
We examined what happens when a published machine learning model for post-transfer mortality, SafeNET, is deployed in an independent health system with substantial data fragmentation and missingness and assessed the extent to which local model development can compensate for cross-system data limitations.
Derivation Cohort:
SafeNET was evaluated from published specifications using its original 14-variable feature schema and applied without modification. Locally trained gradient-boosting models (Categorical Boosting [CatBoost], Light Gradient Boosting, Extreme Gradient Boosting) were developed using a stratified 70/30 train-test split with five-fold cross-validation, class weighting, and random under-sampling to address class imbalance (~1:27).
Validation Cohort:
We included 14,728 adult patients (≥ 18 yr) undergoing IHT at a multihospital academic health system from January 1, 2023, to December 31, 2023. Functional status variables were missing in 85% of encounters, limiting direct implementation of the original feature set.
Prediction Model:
Discrimination, classification metrics (precision, recall, F1 score), and calibration (Brier score) were compared across SafeNET and locally trained models on the held-out test set.
Results:
SafeNET achieved performance that closely approximated its development-site results (area under the receiver operating characteristic curve [AUC-ROC], 0.87; F1, 0.73; 95% CI, 0.69-0.77), indicating strong external consistency. Locally trained CatBoost models outperformed SafeNET (AUC-ROC, 0.92; F1, 0.74; 95% CI, 0.72-0.78) and demonstrated improved calibration (Brier 0.04 vs. 0.055). Key variables, including functional status, were missing in 85% of encounters, limiting direct implementation of the original feature set. Although recall exceeded 0.70 across most models, precision remained below 080 for several configurations, indicating important trade-offs relevant to clinical deployment.
Conclusions:
SafeNET's performance was reproducible in an external health system, supporting its portability. However, local models trained on system-specific electronic health record patterns achieved markedly higher accuracy and calibration. Real-world replication revealed major data-availability barriers and highlighted the tension between acceptable recall and suboptimal precision in critical-care prognostic applications.