Related Experiment Videos
Development of an interpretable machine learning model for predicting in-hospital mortality in ICU patients with
Tianyu Zhao1,2, Kexin Wen1,3, Xumin Han4
1Department of Emergency Medicine and Critical Care Medicine (Ward 2), Hebei General Hospital, Shijiazhuang, Hebei, China.
Background:
Sepsis ranks among the primary causes of mortality in intensive care units (ICUs), often resulting in multiple organ dysfunction due to dysregulated systemic inflammation. Timely recognition of patients at high risk is essential, yet existing clinical scoring systems show limited predictive performance. We aimed to design and evaluate an interpretable machine learning model to predict the risk of in-hospital mortality in individuals with sepsis.
Methods:
This retrospective study screened 503 sepsis patients who were admitted to the ICU at Hebei Provincial People's Hospital from January 2022 through January 2025. After excluding patients with ICU stay less than 24 h, incomplete clinical data, or pregnancy, 431 patients were included. The dataset was randomly partitioned into a training set (n = 301) and a validation set (n = 130). Eight machine learning algorithms-including logistic regression, naive Bayes, support vector machine, single-layer neural network, random forest (RF), LightGBM, K-nearest neighbors, and decision tree-were trained and evaluated. Model performance was assessed by ROC curves (AUC), accuracy, F1 score, sensitivity, specificity, Brier score, expected calibration error, and decision curve analysis. The best-performing model was interpreted using SHAP values.
Results:
Among 431 patients, 198 (45.9%) died during hospitalization. Feature selection using the Boruta algorithm identified 12 key predictors. The RF model achieved the most balanced overall performance, with an area under the ROC curve of 0.804 (95% CI, 0.730-0.879) in the validation set, accuracy of 0.785, F1 score of 0.778, sensitivity of 0.817, and specificity of 0.757. SHAP analysis identified lactate dehydrogenase and urea as the most influential predictors, followed by age, APACHE II score, and procalcitonin. An interactive web application was developed as a research prototype to demonstrate the feasibility of model deployment.
Conclusions:
We developed and internally validated an interpretable RF-based machine learning model for predicting in-hospital mortality in patients with sepsis. The model showed acceptable discrimination, reasonable calibration, and potential clinical net benefit in internal validation. SHAP-based interpretation provided insight into key predictive factors, while the web-based calculator demonstrated the feasibility of model deployment. However, external and prospective validation is required before its clinical utility and generalizability can be established.