Interpretable Machine Learning for Predicting Radiation-Induced Hypothyroidism in Head and Neck Cancer: A Dual-Center
Wenhui Li1, Ying Zhang1, Yangyang Ji2
1Department of Radiation Oncology, The Second Affiliated Hospital of Harbin Medical University, Harbin 150086, China.
Abstract:
Background/Objectives: Radiation-induced hypothyroidism (RIHT) is a frequent complication in head and neck cancer (HNC) patients following radiotherapy, significantly impacting quality of life. This study aimed to develop and validate an interpretable machine learning model for predicting RIHT, with a focus on providing individualized risk assessment. Methods: The study included a development cohort (n = 256) from the Second Affiliated Hospital of Harbin Medical University and an independent external validation cohort (n = 296) from the First Affiliated Hospital of Kunming Medical University. Using 18 predictive features, six machine learning algorithms-Random Forest, Support Vector Machine, Multi-Layer Perceptron, Logistic Regression, XGBoost, and LightGBM-were constructed. Model performance was evaluated using AUC, balanced accuracy, F1 score, PR-AUC, Brier score, and Brier skill score. SHapley Additive exPlanations (SHAP) was applied for model interpretability. Results: In internal validation, LightGBM achieved the highest performance (AUC = 0.872, balanced accuracy = 0.703, PR-AUC = 0.712, and Brier skill score = 0.371), outperforming logistic regression. In external validation, Random Forest achieved the highest AUC (0.769), while XGBoost and LightGBM achieved AUCs of 0.721 and 0.716, respectively. SHAP analysis identified pre-treatment thyroid volume, T-stage, tumor site, mean thyroid dose (Dmean), and V50 as the most influential features for model prediction. Notably, dose-volume parameters (V30-V60) were highly correlated (r = 0.75-0.98); therefore, they should be interpreted as a cluster rather than as independent predictors. We caution against overinterpreting individual Vx thresholds as rigid clinical cutoffs. Conclusions: Machine learning models, particularly tree-based ensemble methods (LightGBM and XGBoost), demonstrate good predictive performance for RIHT. Combined with SHAP analysis, these models provide transparent, individualized risk assessments. However, the dose-response relationship for RIHT is continuous, and no single dose threshold should be used as a rigid clinical cutoff. Model code is available from the corresponding author upon reasonable request. Future prospective studies with time-to-event analysis are needed to further validate clinical utility.
