Early mortality risk prediction in critically ill patients with liver cancer using interpretable machine learning: a
Qin Li1, Bo Zhang2, Lei Zhang3
1Infectious Disease Department, Mianyang 404 Hospital, Mianyang, China.
Objective:
This study aimed to develop and internally evaluate interpretable machine learning models for predicting 28-day all-cause mortality among patients with liver cancer admitted to the intensive care unit (ICU).
Methods:
A total of 507 ICU patients with liver cancer were identified from the Medical Information Mart for Intensive Care IV database. Patients were randomly divided into a training set and a test set at a ratio of 7:3 using stratified sampling. Candidate predictors were restricted to variables available at ICU admission or within the first 24 h after ICU admission. Feature selection was performed using clinical screening, collinearity assessment, and least absolute shrinkage and selection operator regression. Four models were developed, including logistic regression (LR), least absolute shrinkage and selection operator logistic regression (LASSO logistic regression), random forest (RF), and extreme gradient boosting (XGBoost). Model performance was evaluated using AUROC, AUPRC, Brier score, sensitivity, specificity, positive predictive value, negative predictive value, and F1 score. SHapley Additive exPlanations were used to interpret the best-performing model. A simplified model based on key predictors was also constructed.
Results:
Among the 507 included patients, 150 patients died within 28 days after ICU admission, corresponding to a mortality rate of 29.6%. The random forest model achieved the best performance, with an AUROC of 0.974, AUPRC of 0.950, Brier score of 0.065, sensitivity of 84.4%, specificity of 98.1%, and F1 score of 0.894. It outperformed logistic regression, LASSO logistic regression, XGBoost, and conventional severity scores, including SOFA and MELD. SHAP analysis identified APS III, Charlson Comorbidity Index, MELD, blood urea nitrogen, and glucose as important predictors. The simplified 12-variable model retained good discrimination, with an AUROC of 0.890.
Conclusion:
An interpretable random forest model showed high internal predictive performance for 28-day mortality among ICU patients with liver cancer. External and prospective validation is required before clinical implementation.

