Related Experiment Video
Updated: May 20, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
False hope of a single generalisable AI sepsis prediction model: bias and proposed mitigation strategies for
Rudolf Schnetler1,2, Anton van der Vegt3, Vikrant R Kalke4
1Townsville Institute of Health Research and Innovation, Townsville Hospital and Health Service, Townsville, Queensland, Australia.
Objective:
To identify bias in using a single machine learning (ML) sepsis prediction model across multiple hospitals and care locations; evaluate the impact of six different bias mitigation strategies and propose a generic modelling approach for developing best-performing models.
Methods:
We developed a baseline ML model to predict sepsis using retrospective data on patients in emergency departments (EDs) and wards across nine hospitals. We set model sensitivity at 70% and determined the number of alerts required to be evaluated (number needed to evaluate (NNE), 95% CI) for each case of true sepsis and the number of hours between the first alert and timestamped outcomes meeting sepsis-3 reference criteria (HTS3). Six bias mitigation models were compared with the baseline model for impact on NNE and HTS3.
Results:
Across 969 292 admissions, mean NNE for the baseline model was significantly lower for EDs (6.1 patients, 95% CI 6 to 6.2) than for wards (7.5 patients, 95% CI 7.4 to 7.5). Across all sites, median HTS3 was 20 hours (20-21) for wards vs 5 (5-5) for EDs. Bias mitigation models significantly impacted NNE but not HTS3. Compared with the baseline model, the best-performing models for NNE with reduced interhospital variance were those trained separately on data from ED patients or from ward patients across all sites. These models generated the lowest NNE results for all care locations in seven of nine hospitals.
Conclusions:
Implementing a single sepsis prediction model across all sites and care locations within multihospital systems may be unacceptable given large variances in NNE across multiple sites. Bias mitigation methods can identify models demonstrating improved performance across most sites in reducing alert burden but with no impact on the length of the prediction window.
Insights
A single machine learning (ML) sepsis prediction model shows significant bias across hospitals and care settings. Tailored models, trained separately for emergency departments (EDs) and wards, reduce alert burden without affecting prediction time.
Area of Science:
- Medical Informatics
- Clinical Decision Support
- Machine Learning in Healthcare
Background:
- Sepsis prediction models are crucial for timely intervention.
- A single machine learning (ML) model may exhibit performance disparities across diverse hospital settings and care locations.
- Identifying and mitigating bias in ML models is essential for equitable and effective healthcare delivery.
Purpose of the Study:
- To identify bias in a single ML sepsis prediction model across multiple hospitals and care settings.
- To evaluate the impact of six bias mitigation strategies on model performance.
- To propose a generic modeling approach for developing high-performing sepsis prediction models.
Main Methods:
- A baseline ML model was developed using retrospective data from 969,292 patient admissions across nine hospitals.
- Model performance was assessed using the number needed to evaluate (NNE) and hours to sepsis prediction (HTS3).
- Six bias mitigation strategies were compared against the baseline model for their impact on NNE and HTS3.
Main Results:
- The baseline ML model demonstrated significant differences in NNE between emergency departments (EDs) and wards.
- Bias mitigation models improved NNE but did not significantly alter HTS3.
- Models trained separately on ED or ward data achieved the lowest NNE with reduced interhospital variance in seven out of nine hospitals.
Conclusions:
- A single, universal ML sepsis prediction model may be unsuitable for multihospital systems due to performance variations.
- Bias mitigation techniques can enhance model performance in reducing alert burden across sites.
- Current bias mitigation strategies do not extend the prediction window for sepsis detection.

