Related Experiment Videos
Predictive performance of machine learning models for bloodstream infection by validation type: A systematic review
Heejeong Kwon1, Jaewoong Kim2, Hyunok Yun3
1School of Medicine, Hanyang University College of Medicine, Seoul, Korea.
Purpose:
Bloodstream infection (BSI) is a life-threatening condition for which culture-based diagnosis is slow, motivating interest in machine learning (ML) models that predict BSI from routine clinical data. To date, this evidence has been summarized only narratively, without quantitative synthesis or formal appraisal of methodological quality. We performed a systematic review and meta-analysis of ML-based BSI prediction models to quantify their discriminative performance and appraise the certainty of the evidence.
Methods:
We searched five databases for English-language studies published between 2013 and 2025. The area under the receiver operating characteristic curve (AUROC) was the primary discrimination measure and was pooled using random-effects models. Risk of bias was assessed with the Prediction model Risk Of Bias Assessment Tool (PROBAST), and certainty of evidence was assessed with GRADE.
Results:
Twenty-six studies were included, of which 24 (28 estimates) contributed to the meta-analysis. The pooled AUROC was 0.82 (95 % confidence interval, 0.77-0.86), with extreme heterogeneity (I2 = 99.9 %) and a wide 95 % prediction interval (0.46-0.96). Discrimination declined as validation became more stringent, from internal validation (0.86) to external/temporal validation (0.77; P = 0.005), indicating a generalizability gradient, whereas performance did not differ significantly by clinical setting (P = 0.52) or target organism (P = 0.53). Most studies (96 %) were at high risk of bias in the PROBAST analysis domain, and the overall certainty of evidence was very low.
Conclusion:
ML models showed promising discrimination for BSI prediction (pooled AUROC 0.82). However, the prediction interval was wide (0.46-0.96) and performance declined under external validation, indicating limited generalizability; moreover, the current evidence is of very low certainty. Rigorous external validation, calibration and decision-analytic reporting, and adherence to prediction-model reporting standards are needed before clinical implementation.