Related Experiment Videos
Interpretable admission-stage machine learning for early Sepsis-3 identification in the emergency department
Fulden Cantaş Türkiş1, Buğra Varol1
1Division of Biostatistics, Faculty of Medicine, Muğla Sıtkı Koçman University, Muğla, Türkiye.
Background/Objectives:
Artificial intelligence-based clinical decision support systems have emerged as promising tools for improving early sepsis recognition in emergency departments. This study aimed to develop and evaluate interpretable machine learning models for identifying Sepsis-3-defined sepsis using routinely available admission-stage variables and to compare their performance with conventional biomarkers and the Systemic Inflammatory Response Syndrome (SIRS).
Methods:
This secondary analysis used a publicly available prospective cohort of 1,572 adult emergency department patients with suspected community-onset sepsis. The outcome was Sepsis-3-defined sepsis. Thirteen demographic, physiological, and laboratory admission variables were evaluated. Following a stratified 70:30 train-test split, missing data were handled using multivariate imputation by chained equations, and feature selection was performed using least absolute shrinkage and selection operator logistic regression. Penalized logistic regression (LR), Explainable Boosting Machine (EBM), Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM) models were developed using five-fold cross-validation. Performance was evaluated by discrimination, classification, calibration, decision curve analysis, and explainability.
Results:
Among 1,572 adult patients, 560 met the Sepsis-3 criteria. In the internal holdout test set, AUROC values were 0.753 (LR), 0.763 (XGBoost), 0.767 (LightGBM), and 0.766 (EBM). Conventional biomarkers and SIRS showed lower discrimination (AUROC 0.599-0.693), whereas a biomarker-only logistic regression model achieved an AUROC of 0.710. No significant AUROC differences were observed among the machine learning models. LightGBM achieved the numerically highest AUROC and accuracy, whereas EBM achieved the highest sensitivity (0.845). Age, procalcitonin, oxygen saturation, respiratory rate, and neutrophil-to-lymphocyte ratio were the most influential predictors.
Conclusions:
Interpretable machine learning models based on routinely available emergency department admission variables provided moderate discrimination, clinically relevant sensitivity, and transparent predictor effects for identifying Sepsis-3-defined sepsis. Compared with conventional biomarkers, SIRS, and a biomarker-only logistic regression model, the machine learning models generally showed higher discrimination, although the magnitude and statistical significance of these differences varied across comparators. External validation is required before clinical implementation.