Related Experiment Videos
Development and Validation of a Machine Learning-Based Prediction Model for Postoperative Delirium After Glioma
1Department of Neurosurgery, The First Affiliated Hospital of Naval Medical University (Changhai Hospital), Shanghai 200433, China.
Background:
Postoperative delirium after glioma surgery is difficult to predict because risk information evolves throughout the perioperative period. We developed and internally validated a temporally structured machine-learning framework for early risk stratification.
Methods:
This single-centre retrospective study included adults who underwent resection of pathologically confirmed glioma between December 2021 and December 2025. Predictors were classified according to their temporal availability. A perioperative reference model included variables available by the end of surgery, whereas the primary fixed 24-h postoperative landmark model additionally incorporated ICU admission status ascertained by the 24-h landmark. Feature selection was performed in the training cohort using clinical review, correlation analysis, and LASSO regression. Six machine-learning algorithms were evaluated for discrimination, calibration, classification performance, clinical utility using decision-curve analysis, and model interpretability using SHAP.
Results:
Among 316 eligible patients, 221 were assigned to the training cohort and 95 to the internal validation cohort. Postoperative delirium occurred in 44 and 19 patients, respectively. The perioperative model retained six predictors: age, ASA III-IV status, tumour size, neutrophil-to-lymphocyte ratio, albumin, and intraoperative blood loss. ICU admission was interpreted as an early postoperative predictive marker rather than a causal risk factor. XGBoost achieved the highest numerical validation AUC (0.895) and the lowest Brier score (0.134), with an accuracy of 0.832, sensitivity of 0.842, specificity of 0.829, and F1-score of 0.667. Paired DeLong testing showed that XGBoost had significantly higher discrimination than logistic regression, whereas differences from the other machine-learning models were not statistically significant. SHAP analysis identified clinically interpretable contributions from perioperative and early postoperative predictors.
Conclusion:
A temporally structured machine-learning framework showed promising internal validation performance for early POD risk stratification after glioma surgery. External multicentre validation and prospective evaluation are required before clinical implementation.