False hope of a single generalisable AI sepsis prediction model: bias and proposed mitigation strategies for

Rudolf Schnetler1,2, Anton van der Vegt3, Vikrant R Kalke4

  • 1Townsville Institute of Health Research and Innovation, Townsville Hospital and Health Service, Townsville, Queensland, Australia.

BMJ Quality & Safety
|March 26, 2025
PubMed
Abstract

Insights

A single machine learning (ML) sepsis prediction model shows significant bias across hospitals and care settings. Tailored models, trained separately for emergency departments (EDs) and wards, reduce alert burden without affecting prediction time.

Area of Science:

  • Medical Informatics
  • Clinical Decision Support
  • Machine Learning in Healthcare

Background:

  • Sepsis prediction models are crucial for timely intervention.
  • A single machine learning (ML) model may exhibit performance disparities across diverse hospital settings and care locations.
  • Identifying and mitigating bias in ML models is essential for equitable and effective healthcare delivery.

Purpose of the Study:

  • To identify bias in a single ML sepsis prediction model across multiple hospitals and care settings.
  • To evaluate the impact of six bias mitigation strategies on model performance.
  • To propose a generic modeling approach for developing high-performing sepsis prediction models.

Main Methods:

  • A baseline ML model was developed using retrospective data from 969,292 patient admissions across nine hospitals.
  • Model performance was assessed using the number needed to evaluate (NNE) and hours to sepsis prediction (HTS3).
  • Six bias mitigation strategies were compared against the baseline model for their impact on NNE and HTS3.

Main Results:

  • The baseline ML model demonstrated significant differences in NNE between emergency departments (EDs) and wards.
  • Bias mitigation models improved NNE but did not significantly alter HTS3.
  • Models trained separately on ED or ward data achieved the lowest NNE with reduced interhospital variance in seven out of nine hospitals.

Conclusions:

  • A single, universal ML sepsis prediction model may be unsuitable for multihospital systems due to performance variations.
  • Bias mitigation techniques can enhance model performance in reducing alert burden across sites.
  • Current bias mitigation strategies do not extend the prediction window for sepsis detection.

Related Concept Videos