Evaluating deep learning sepsis prediction models in ICUs under distribution shift: a multi-centre retrospective
Fanny Tranchellini1, Youssef Farag1, Catherine Jutzeler1,2
1Department of Health Sciences and Technology, ETH Zurich, Zurich, Switzerland.
NPJ Digital Medicine
|March 3, 2026
Summary
Sepsis prediction models struggle with generalizability due to data shifts. New strategies like retraining and domain adaptation outperform traditional fine-tuning, especially with varying data availability.
Area of Science:
- Critical Care Medicine
- Artificial Intelligence in Healthcare
- Biomedical Informatics
Background:
- Sepsis prediction models trained on intensive care unit (ICU) data often exhibit poor generalization on external datasets due to distribution shifts.
- Existing research primarily focuses on direct model deployment or conventional transfer learning, such as fine-tuning, with limited exploration of alternative strategies.
Purpose of the Study:
- To systematically evaluate and compare various model deployment strategies for sepsis prediction under distribution shift.
- To quantify data shifts across harmonized adult ICU cohorts and assess the performance of different strategies across varying target data availability.
Main Methods:
- Quantified distribution shifts across three harmonized adult ICU cohorts (HiRID, MIMIC-IV, eICU) involving 216,536 patient stays.
- Compared five deployment strategies: generalization, fine-tuning/retraining, target training, supervised domain adaptation (DA), and fusion-training.
- Evaluated strategies across multiple deep learning architectures and four target-data regimes (none, small, medium, large).
Main Results:
- Fine-tuning consistently underperformed compared to other methods.
- Retraining and fusion-training demonstrated superior performance in small and large target data regimes, respectively.
- Supervised domain adaptation (DA) provided the most stable performance gains, particularly in the medium target data regime, improving AUROC and normalized AUPRC.
Conclusions:
- Routine fine-tuning is not optimal for sepsis prediction models facing distribution shifts.
- The choice of deployment strategy should be guided by the availability of target data and specific operational contexts.
- Domain adaptation and retraining/fusion offer more robust and effective approaches for generalizing sepsis prediction models.

