Related Experiment Videos
Development and validation of an interpretable machine learning model for predicting postoperative fever after
Xiaofei Lu1,2, Chunping Yu2, Zhiyong Ding2
1Beijing Chaoyang Hospital, Capital Medical University, Beijing, China.
Background:
This study aimed to develop and validate an interpretable machine learning model to predict postoperative fever (POF) after flexible ureteroscopic lithotripsy (fURL).
Objective:
This study adopted a single-center retrospective temporal cohort design. Strictly adhering to predefined inclusion and exclusion criteria, patients with upper urinary tract stones (maximum diameter < 3 cm) who underwent flexible ureteroscopic lithotripsy (fURL) from January 2019 to December 2024 were enrolled as the derivation cohort, while consecutive cases treated between January and December 2025 were recruited to construct an independent temporal validation cohort. Least absolute shrinkage and selection operator (LASSO) regression filtered optimal predictors for postoperative fever (POF) risk modeling. Six machine learning (ML) algorithms included logistic regression (LR), random forest (RF), multi-layer perceptron (MLP), support vector machine (SVM), XGBoost, and LightGBM were built using combined clinical and imaging features to predict fURL-related POF. Model performance was assessed via multiple metrics including AUC and F1-score. The SHAP algorithm was employed to perform multi-dimensional interpretability analysis on the optimal model, thereby quantifying the contribution magnitude and predictive mechanism of each feature. Furthermore, the temporal validation cohort was used to evaluate the generalizability and extrapolation stability of the established model.
Results:
Among the six machine learning algorithms, The LightGBM model yielded the optimal discriminative metrics in this model comparison, including the maximum AUC (0.896, 95% CI: 0.853-0.940) and PR-AUC (0.787), as well as the highest specificity (0.967) and accuracy (0.891). DeLong pairwise comparisons demonstrated that LightGBM was statistically superior to four competing models; however, the AUC difference between LightGBM and XGBoost did not reach statistical significance. The SHAP analysis further clarifies the contribution of each variable in the prediction process. Temporal validation revealed acceptable temporal generalization performance of the LightGBM predictive model, with an AUC of 0.723 (95% CI: 0.648-0.799).
Conclusion:
This study employed an interpretable machine learning model to effectively predict POF following fURL. This model provides a reference for surgical risk stratification, which helps establish personalized postoperative monitoring strategies.