Related Experiment Video
Updated: Sep 16, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Uncertainty estimation for reliable neural network-based radiotherapy dose modelling using Monte Carlo dropout, mean
Moritz Schneider1, Tabea Eberhardt1, Cihan Gani2
1Section for Biomedical Physics, Department of Radiation Oncology, University Hospital Tübingen, Tübingen, Germany.
Background And Purpose:
Neural networks promise fast dose modelling with high accuracy for challenging situations like magnetic resonance imaging (MRI)-guided radiotherapy. As they are data-driven, failure can occur and early identification of erroneous dose calculations is required.In this study, we implemented and evaluated three uncertainty estimation techniques to assess whether they can indicate increased prediction error.
Materials And Methods:
Using a dataset of 6713 radiotherapy segments from 130 1.5 T MRI linear accelerator plans and a 3D UNet for dose modelling, three techniques were implemented to assess uncertainty: Monte Carlo dropout (MCD), mean variance estimation (MVE) and a deep ensemble (DE).All methods were evaluated regarding calibration using the expected normalized calibration error (ENCE) and correlation to mean absolute error (MAE).Cumulative failure rates were calculated across uncertainty thresholds, with a 3 mm/3% gamma passing rate < 95% defining failure.
Results:
After calibration, all three methods showed ENCE values of 0.14/0.05/0.10 for MCD/MVE/DE. A strong relationship between mean uncertainty and MAE was observed, reflected by high Spearman correlation coefficients (MCD: ρ = 0.64, MVE: ρ = 0.76, DE: ρ = 0.67).Failure rates increased with mean uncertainty across all methods, enabling thresholds targeting a 10% failure rate. On independent evaluation, 44%-77% of segments were accepted, with observed failure rates among accepted segments of 7.2%-9.5%.
Conclusions:
All methods provided meaningful uncertainty estimates clearly associated with prediction error. MVE showed the lowest ENCE and strongest correlation with MAE while requiring the least computation. DE most closely matched the target failure rate while flagging the fewest segments.