Related Experiment Video
Updated: Aug 10, 2026

Proton Therapy Delivery and Its Clinical Application in Select Solid Tumor Malignancies
Published on: February 6, 2019
Patient-Specific Deep Reinforcement Learning for Proton Beam Delivery Under Inter-Phase Variations
Mélanie Ghislain1, Estelle Loÿen1, Antoine Aspeel2
1UCLouvain (ICTEAM), Place du Levant 3, Louvain-la-Neuve, 1348, Belgium.
Purpose:
Proton therapy is challenged by tumor motion, particularly for lung tumors affected by respiratory-induced motion. Conventional planning strategies compensate for this motion by introducing safety margins or by using robust optimization, increasing irradiation of surrounding healthy tissues. Real-time plan adaptation during delivery represents a promising alternative to mitigate intrafractional motion effects.
Materials And Methods:
We propose a patient-specific deep Reinforcement Learning (RL)-based control framework for proton pencil beam scanning, formulated as a first step toward a real-time plan adaptation problem under respiratory motion. RL agents are trained independently on 3 patients on their mid-position planning CT to sequentially control beam position and spot delivery. At inference, the learned policy is executed on the respiratory phases of each patient's 4DCT without any online re-training or model parameters adaptation. The agent is provided with 2D observations encoding target geometry, beam position, and prior spot delivery. The approach is evaluated against conventional static gross tumor volume-based (GTV-based) and internal target volume-based (ITV-based) planning strategies.
Results:
The agents improve target coverage under respiratory motion by exploiting new information in the observations during delivery compared with static gross tumor volume-based plans, with an average gain of 4.54 Gy in over the whole treatment for Patient 1. Compared with ITV-based plans, the RL-based approach generally reduces dose exposure to organ-at-risk, with an average reduction of 0.42, 4.12, and 3.14 Gy in D for the 3 patients, respectively, and a decrease of 1.41 Gy in D for Patient 3, who has the largest motion amplitude.
Conclusion:
This study should be interpreted as a proof-of-concept and highlights the potential of RL-based control strategies for proton therapy delivery under intrafractional motion. While the current framework relies on static training, it establishes a foundation for future extensions toward fully dynamic and adaptive treatments.

