Integrating patient-reported outcomes into explainable machine learning models for long-term survival prediction in
Miguel Ángel Gómez-Luque1, Guillermo Lendínez-Cano1, Belén Carrero-García1
1Urology and Nephrology Department, Biomedical Institute of Seville (IBIS), University Hospital Virgen del Rocío, 41013 Seville, Spain.
Background:
Long-term survival (LTS, overall survival ≥48 months) in metastatic renal cell carcinoma (mRCC) remains an incompletely understood phenotype. Established prognostic models such as the IMDC criteria do not incorporate patient-reported outcomes (PROs) and were not designed to identify individual patients with potential for long-term disease control. Machine learning (ML) methods may improve LTS classification by capturing non-linear interactions across heterogeneous variable sets, while explainability frameworks such as SHAP allow clinical interpretation of model predictions.
Methods:
In a retrospectively collected, single-centre cohort of 71 mRCC patients treated with first-line VEGF-targeted TKI (LTS: n=25, 35.2%; nLTS: n=46, 64.8%), five supervised classification algorithms were developed and evaluated: K-Nearest Neighbors (KNN, k=10), Support Vector Machine with Gaussian kernel (SVMGAM), Random Under-Sampling Boosting (RUSBoost), Linear Discriminant Analysis (RLI), and Generalized Additive Model (GAM). The feature set comprised 54 baseline variables encompassing demographics, comorbidities, laboratory parameters, tumour and metastatic disease characteristics, and all 19 individual items, three subscale scores (FKSI-DRS-P, FKSI-TSE, FKSI-F/WB), and total score of the NCCN-FKSI-19 questionnaire. Model performance was assessed by stratified 5-fold cross-validation using the area under the ROC curve (AUC), sensitivity, specificity, and F-score, each with bootstrap 95% confidence intervals. Robustness was further evaluated by repeated cross-validation, LASSO-penalised modelling with nested cross-validation, and decision-curve analysis against IMDC risk. SHAP values were computed for the primary model (RUSBoost) across all patients to identify the most important features.
Results:
RUSBoost achieved the highest discrimination (AUC 0.863, 95% CI 0.766-0.945; sensitivity 0.760, specificity 0.804, F-score 0.717), followed by SVMGAM (AUC 0.830) and KNN10 (AUC 0.816). Cross-fold AUC varied between algorithms, with SVMGAM the most stable and RUSBoost the most variable. Under repeated cross-validation, mean discrimination was more conservative (AUC 0.765 ± 0.061), and a LASSO-penalised model with feature selection nested within the cross-validation achieved comparable discrimination (0.790 ± 0.034); the model showed higher net benefit than IMDC risk stratification across clinically relevant thresholds. SHAP analysis identified haemoglobin, the pain item GP4, dyslipidaemia, and nephrectomy status as the most important features. GP4 was the only FKSI-19 item among the leading predictors and ranked second overall across the 54-variable set and was selected in 100% of LASSO resamples, suggesting that individual PRO items carry greater discriminative weight than their aggregate scores.
Conclusions:
In this retrospective single-centre cohort, ML algorithms, particularly ensemble classifiers addressing class imbalance, classified LTS in TKI-treated mRCC using exclusively baseline data. SHAP-based explainability suggests that the predictive signal within the FKSI-19 questionnaire is concentrated at the individual item level. These findings should be regarded as hypothesis-generating and require prospective external validation before PRO item-level data can be integrated into clinical prognostic workflows.
Related Concept Videos
Kaplan-Meier Approach
Cancer Survival Analysis
