Improving deep-learning IMRT plan quality for lymphoma using deliverability-aware training
Denis Kutnár1,2, Ivan Richter Vogelius2,3, Hristo Atanasov Georgiev1,2
1Department of Computer Science, University of Copenhagen, Copenhagen Ø, Denmark.
Background:
Deep learning (DL) methods have shown strong performance in predicting dose distributions from patient anatomy and target structures; however, accurate dose prediction does not guarantee a deliverable plan after conversion into machine parameters and dose recalculation in the treatment planning system (TPS), motivating training objectives that explicitly couple prediction to deliverability. Lymphoma presents substantial clinical heterogeneity in target location, extent, and prescription/fractionation, so fast automated methods that produce deliverable plans are particularly valuable.
Purpose:
To investigate whether a deliverability-aware training framework improves the end-to-end deliverability of DL-generated intensity-modulated radiotherapy (IMRT) plans for lymphoma.
Methods:
A retrospective single-institution dataset of 619 plans (571 patients; 2009-2021) was used. Clinical volumetric modulated arc therapy (VMAT) plans were converted to standardized 15-field coplanar sliding-window IMRT reference plans in Eclipse using automated scripting. Two DL approaches were compared: a baseline dose-only objective model (DL-DO) and a multi-objective model (DL-MO) trained with dose supervision, fluence supervision, and a dose-fluence-dose cycle-consistency constraint implemented via fixed inverse and forward networks. Predicted fluence maps were imported into the TPS, leaf-sequenced, and recalculated with AcurosXB without further optimization; plan quality was evaluated on a held-out test set (n = 61) using planning target volume (PTV) D98, D2, and Dmax with paired Wilcoxon testing and Holm correction.
Results:
DL-MO produced significantly better machine-deliverable plans than DL-DO, improving coverage and reducing hotspots (median paired differences: +12.31 percentage points (pp) in D98, -3.43 pp in D2, and -3.18 pp in Dmax; Holm-corrected p < 0.0001 for all). Mean ± SD PTV metrics were D98 = 90.10 ± 3.74% vs. 78.59 ± 5.92%, D2 = 106.28 ± 1.13% vs. 109.65 ± 2.40%, and Dmax = 110.32 ± 2.20% vs. 113.27 ± 3.03% (DL-MO vs. DL-DO). DL-MO approached the TPS reference IMRT plans (D98 = 91.62 ± 3.45%) with residual hotspot elevation.
Conclusions:
Cycle-consistency training improves the end-to-end deliverability of DL-generated IMRT plans for heterogeneous lymphoma, yielding substantially better TPS-recalculated, machine-deliverable plans than conventional dose-only supervision while maintaining sub-second fluence-map generation. However, statistically significant differences from the TPS reference remained for D98, D2, and Dmax.

