Development and validation of an interpretable machine-learning model for predicting treatment failure in severe
Qiang Shi1, Yun Liu1, Jiahui Shen2
1Department of Emergency, Suzhou Ninth People's Hospital, Suzhou Ninth Hospital Affiliated to Soochow University, Suzhou, China.
Objective:
This study aims to develop and validate machine learning (ML) models for predicting treatment failure in trauma patients using comprehensive clinical and laboratory variables, and to identify key prognostic features.
Methods:
A retrospective cohort of 318 trauma patients was included. We included 44 characteristics, and the primary outcome was treatment failure at hospital discharge, defined as in-hospital death, an unimproved or worsened discharge status, discharge against medical advice or withdrawal of active treatment because of critical illness, or a GOS score of 1-3 in patients with concomitant traumatic brain injury. The dataset was randomly divided into a training set (70%) and a test set (30%). Features were selected via least absolute shrinkage and selection operator (LASSO) regression in the training set. 5 ML models-Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), k-Nearest Neighbors (KNN), and Extreme Gradient Boosting (XGBoost)-were trained and evaluated. The optimal model was interpreted using SHapley Additive exPlanations (SHAP) analysis.
Results:
LASSO regression selected 13 candidate predictors for model development in the training set. The RF model demonstrated the best performance in the test set, with an AUC of 0.967 (95% CI: 0.936-0.997), sensitivity of 1.000, specificity of 0.855, and F1-score of 0.776. SHAP analysis identified Glasgow Coma Scale (GCS) score as the most influential predictor, followed by Multiple Organ Dysfunction Syndrome (MODS), Acute PHysiology and Chronic Health Evaluation (APACHE II) score, Injury Severity Score (ISS), and creatine kinase (CK). Higher APACHE II and ISS scores were positively associated with treatment failure, while higher GCS, absence of MODS and CK levels correlated with reduced risk.
Conclusion:
Utilizing multiple trauma severity scores, laboratory parameters, and machine learning algorithms, we developed a predictive model to identify trauma patients at risk of treatment failure within the first 24 h of admission. Among the algorithms evaluated, RF model demonstrated superior internal discriminative performance in our single-center cohort. SHAP interpretability analysis reveals the core prognostic value of GCS, MODS, APACHE II, ISS, and CK, providing a potential transparent auxiliary tool for clinical risk stratification. However, its generalization performance still needs multi center external validation.
