An evaluation of uncertainty quantification methods and measures for deep learning outcome prediction models in head
Daniel C MacRae1, Luuk van der Hoek1, Joëlle E van Aalst1
1Department of Radiation Oncology, University Medical Center Groningen, University of Groningen, Groningen, the Netherlands.
Background And Purpose:
Deep learning (DL) outcome prediction models show promise in radiotherapy but face limited clinical adoption due to concerns about prediction reliability. Although uncertainty quantification (UQ) can provide confidence estimates alongside predictions, there is currently little consensus on appropriate UQ approaches for DL-based outcome prediction. This study therefore evaluates and compares different UQ approaches for normal tissue complication probability (NTCP) and tumour control probability (TCP) models in head and neck cancer.
Materials And Methods:
Four published DL models were reproduced: two NTCP models predicting xerostomia and dysphagia at six months post-treatment, and two TCP models predicting two-year survival and locoregional control. Three UQ methods-Monte Carlo dropout, deep ensembles, and test-time augmentation-and three uncertainty measures-predictive entropy, variance, and mutual information-were evaluated. Models were retrained on development sets (NTCP: 964 patients, TCP: 255 patients), and assessed on independent validation sets (NTCP: 241 patients, TCP: 85 patients) using discriminative performance, calibration metrics, and sparsification analysis.
Results:
Incorporating UQ methods maintained comparable discriminative performance to baseline models across all endpoints (mean AUC: 0.72-0.73 vs 0.72). Deep ensembles and Monte Carlo dropout demonstrated strong calibration between uncertainty values and prediction accuracy, while test-time augmentation showed variable reliability. Entropy and variance consistently correlated with prediction accuracy, whereas mutual information proved unstable.
Conclusions:
Monte Carlo dropout and deep ensembles provide meaningful uncertainty estimates for NTCP and TCP prediction without compromising model performance. These methods show potential for selective prediction workflows where high-confidence predictions guide treatment decisions while uncertain cases are flagged.

