Machine-Learning Based Prognostic Model for Predicting Early Recurrence in HCC Patients After Hepatectomies: An
Heng-Yuan Hsu1, Jiunn-Chang Lin2,3,4, Chun-Wei Huang1,5
1Division of General Surgery, Department of Surgery, New Taipei Municipal Tucheng Hospital, Built and Operated by Chang Gung Medical Foundation, New Taipei, 23652, Taiwan.
Purpose:
Early recurrence within 24 months post-resection remains a primary driver of poor prognosis in hepatocellular carcinoma (HCC). In the absence of standardized adjuvant guidelines, robust postoperative risk stratification is critical. We evaluated explainable machine learning (ML) architectures to optimize risk modeling using readily accessible parameters.
Patients And Methods:
This retrospective, multicenter study analyzed 1,681 HCC patients undergoing curative-intent hepatectomy at Chang Gung institutions (2007-2020) as the training cohort. External validation was conducted using an independent cohort (n = 251) from Mackay Memorial Hospital. Four algorithms-random survival forest, Cox-nnet, LASSO, and extreme gradient boosting (XGBoost)-were trained using 5-fold cross-validation. Missing data were handled via k-nearest neighbors imputation. Discriminative capacity was assessed using the concordance index (C-index), and feature significance was decoded through SHAP values.
Results:
The XGBoost framework yielded optimal discrimination, achieving a high training C-index of 0.98. During independent external validation, the C-index attenuated to a robust 0.72, reflecting expected adjustments for baseline institutional heterogeneities. Multivariable Cox and SHAP analyses consistently identified five pivotal predictors: sex, preoperative treatment, tumor size, satellite lesions, and vascular invasion. The derived nomogram enabled effective patient risk-tiering (p < 0.0001), although absolute recurrence probabilities were systematically overestimated in the external validation cohort.
Conclusion:
While the XGBoost model exhibits expected calibration shifts across disparate cohorts, it provides robust, cross-center discriminative generalizability for categorical risk stratification. Rather than serving as an absolute probability estimator, this explainable model functions as a reliable clinical tool to selectively identify high-risk candidates for intensive imaging surveillance. Geographically and ethnically diverse prospective validation remains required prior to broader clinical deployment.
