Related Experiment Videos
Classification of recurrence status after surgical treatment of chronic subdural hemorrhage - A machine learning
Hussam Hamou1, Julius Kernbach1,2, Hani Ridwan3
1Department of Neurosurgery, RWTH Aachen University Hospital, Aachen, Germany.
Background:
Chronic subdural hematoma (cSDH) recurrence requiring reoperation occurs in 5-33% of cases. Predicting recurrence could enable risk-stratified surveillance, reducing imaging in low-risk patients while maintaining monitoring for high-risk individuals. We evaluated whether machine learning could achieve clinically actionable recurrence prediction using routinely available variables.
Methods:
This retrospective single-center study included 564 consecutive patients undergoing surgical cSDH evacuation (2015-2023), randomly divided into training (75%, n = 422) and test (25%, n = 142) sets. We developed and compared three models, regularized logistic regression, Random Forest, and XGBoost, using 31 predictor variables. Model development and tuning used 10-fold cross-validation on the training set; the best model was evaluated on the held-out test set. The primary outcome was postoperative recurrence requiring reoperation.
Results:
Postoperative recurrence occurred in 170 patients (30.1%). In the training set, XGBoost achieved the highest cross-validated ROC AUC (0.713, SE = 0.024), matching Random Forest and outperforming logistic regression (0.686). Hematoma volume, coagulation parameters, and disease severity markers (ICU admission, GCS) were the most influential predictors, though effect sizes remained modest. On the test set, the final XGBoost model achieved ROC AUC 0.688 (95% CI 0.590-0.772), with satisfactory average calibration but overconfident individual-level risk estimates (calibration slope 0.615). At the clinically relevant 90% sensitivity threshold, specificity was only 30.3%, allowing potential imaging reduction in roughly one-third of non-recurrence patients. Consistency between training and test performance indicated these limitations reflect predictor information content rather than overfitting.
Conclusions:
In this single-center cohort, machine learning models using routinely available clinical and radiographic variables did not achieve clinically actionable risk stratification for cSDH recurrence under internal validation, with discriminative capacity insufficient to identify a low-risk subgroup suitable for de-escalated surveillance. These findings suggest recurrence is driven by factors not captured in standard clinical assessment, supporting uniform or symptom-driven imaging strategies over risk-stratified approaches.