Related Experiment Video
Updated: Apr 8, 2026

A Reproducible Intensive Care Unit-Oriented Endotoxin Model in Rats
Published on: February 20, 2021
Performance of a Sepsis Prediction Model Across Different Sepsis Definitions
Sayon Dutta1,2,3, Reid McMurry4, Michael C Tasi1,3
1Department of Emergency Medicine, Massachusetts General Hospital, Boston.
Importance:
Early detection of sepsis improves clinical outcomes, but the Early Detection of Sepsis Model, version 1 (Epic Systems Corp) has shown poor performance. Comparing sepsis models is complicated by varying outcome definitions and limited generalizability outside the development site.
Objective:
To evaluate the Early Detection of Sepsis Model, version 2 using multiple standard sepsis definitions.
Design, Setting, And Participants:
This diagnostic study included all adult (aged ≥18 years) encounters from emergency departments, inpatient units, intensive care units, and perioperative areas across 9 acute care hospitals between March 17 and August 31, 2024. The locally trained gradient-boosted tree ensemble sepsis model, incorporating patient demographics, vital signs, laboratory results, medication administration, and other clinical features, generated predictions every 15 minutes. Model performance was evaluated against 3 electronically computable sepsis definitions: the Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3); Centers for Medicare & Medicaid Services Severe Sepsis and Septic Shock: Management Bundle (SEP-1); and the Centers for Disease Control and Prevention Adult Sepsis Event (ASE).
Exposure:
Silent deployment of the sepsis model.
Main Outcomes And Measures:
Model performance metrics included discrimination, as measured using the area under the receiver operating characteristic curve (AUROC), and precision, as measured by the area under the precision-recall curve (AUPRC). Test characteristics and potential lead-time warning were evaluated over a range of model score thresholds.
Results:
Among 198 494 patient encounters (median [IQR] age, 55 [36-71] years; 54.8% female) the incidence of sepsis varied by outcome definition. A total of 5832 encounters (2.9%) met the Sepsis-3 definition, 2366 (1.2%) met the SEP-1 definition, and 3881 (2.0%) met the ASE definition. The AUROC and AUPRC were, respectively, 0.89 (95% CI, 0.89-0.90) and 0.24 (95% CI, 0.23-0.25) for Sepsis-3, 0.94 (95% CI, 0.94-0.94) and 0.16 (95% CI, 0.15-0.17) for SEP-1, and 0.85 (95% CI, 0.85-0.86) and 0.11 (95% CI, 0.11-0.12) for ASE. In the overall study population, the Youden top left threshold for the Sepsis-3 outcome definition was 9, which yielded a positive predictive (PPV) of 11.4% (95% CI, 11.1%-11.7%) and median lead time of 3.4 hours (IQR, 0.9-22.3 hours). The SEP-1 had a PPV of 6.8 (95% CI, 6.6-7.1 9.9) and median (IQR) lead time of 4.5 (4.3-23.1) hours, and the ASE had a PPV of 5.9 (95% CI, 5.7-6.0) and median (IQR) lead time of 1.4 (0.5-14.8) hours.
Conclusions And Relevance:
This cohort study found that the locally trained sepsis model showed moderate predictive performance, with discrimination, precision, and lead time varying considerably by the sepsis definition applied. Although the model provided some advance warning before sepsis onset, a high false-positive rate may limit its clinical utility without careful threshold selection and tailored implementation.

