Related Experiment Video
Updated: Sep 27, 2026

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Externally validated explainable machine learning for postoperative recurrence prediction in early-stage NSCLC
Fatih Kemik1, Hayri Kağan Gören2, Bahadır Köylü1
1Department of Medical Oncology, Koç University School of Medicine, Istanbul, Türkiye.
Background:
Postoperative recurrence remains a major challenge in early-stage non-small cell lung cancer (NSCLC), and pathological TNM staging does not fully capture within-stage heterogeneity. We aimed to develop and externally validate an explainable machine learning model for recurrence prediction after curative resection.
Methods:
This retrospective study included 723 patients with stage I-II NSCLC, including 52 recurrences, with independent external validation in a separate cohort (n=50). A Random Forest model using routinely available clinical and pathological variables was developed within a nested cross-validation framework and compared with logistic regression. Performance was evaluated using ROC-AUC, calibration, Decision Curve Analysis, and SHAP-based interpretation.
Results:
The Random Forest achieved a mean internal ROC-AUC of 0.70 versus 0.64 for logistic regression, although no formal paired statistical comparison was performed. In an independent case-control external validation cohort with an artificially balanced outcome distribution, the Random Forest achieved an ROC-AUC of 0.68 (95% CI 0.53-0.83); the balanced design precluded assessment of calibration or absolute risk at natural prevalence. Sigmoid recalibration of pooled out-of-fold predictions yielded an apparent Brier score reduction to 0.063, although post-calibration performance was not independently evaluated. Decision Curve Analysis demonstrated positive net clinical benefit across clinically relevant thresholds. SHAP analysis highlighted tumor size, pathological T stage, and STAS among the principal tumor-related predictors, whereas fold-level analysis showed greater stability for tumor size, tumor necrosis, and STAS. Complementary time-to-event analyses accounting for right censoring included 715 patients and 44 recurrence events.
Conclusion:
An explainable machine learning model based on routinely available variables demonstrated promising but preliminary discrimination in an independent case-control external validation cohort. Larger consecutive cohorts with natural outcome prevalence are required before clinical application.