Related Experiment Video
Updated: Jan 31, 2026

06:19
Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
2.5K
A tandem reinforcement learning framework for localized prostate cancer treatment planning and machine parameter
Nathan Shaffer1, Avinash Reddy Mudireddy2, Joel St-Aubin3
1Department of Biomedical Engineering, University of Iowa, Iowa City, Iowa, USA.
Medical Physics
|January 30, 2026
Summary
A new deep reinforcement learning (RL) algorithm automates Volumetric Modulated Arc Therapy (VMAT) machine parameter optimization for prostate cancer, generating clinically comparable plans faster than traditional methods.
Area of Science:
- Medical Physics
- Radiation Oncology
- Machine Learning
Background:
- Volumetric Modulated Arc Therapy (VMAT) machine parameter optimization (MPO) is computationally intensive.
- Existing machine learning methods often supplement, not replace, traditional optimizers and rely on extensive training data.
- Deep reinforcement learning (RL) offers a novel approach for VMAT MPO by learning optimal strategies through trial-and-error.
Purpose of the Study:
- To develop and validate a deep RL algorithm for automated VMAT MPO.
- To generate clinically comparable prostate cancer treatment plans independent of commercial treatment planning system (TPS) optimizers.
- To ensure generated plans meet machine constraints.
Main Methods:
- A dataset of 100 prostate cancer patients was used for training, validation, and testing (70-10-20 split).
- A Proximal Policy Optimization (PPO) algorithm trained two convolutional neural networks to optimize multi-leaf collimator (MLC) positions and monitor units (MUs).
- The RL networks used dose, structure masks, and machine parameters as inputs, optimizing a dose-volume histogram (DVH)-based reward function.
Main Results:
- The RL algorithm generated VMAT plans in an average of 6.3 ± 4.7 seconds.
- RL-generated plans showed improved bladder and rectum sparing compared to reference plans.
- Statistically significant improvements were observed in PTV D2% and reduced mean rectal dose (Dmean), with all clinical objectives met.
Conclusions:
- A deep RL framework was successfully developed and validated for VMAT MPO.
- The algorithm rapidly produces high-quality VMAT prostate cancer plans comparable to manual optimization.
- This demonstrates RL's potential to fully automate VMAT planning, reducing time while maintaining quality.
Related Concept Videos
Treatment Resistant Cancers
3.8K
Cancer is the second leading cause of death in the United States. A cancer cell is genetically unstable and hence can mutate faster. They can also modify their microenvironment and escape immune surveillance. The difficulties in treating cancer are further compounded by the emergence of rapid resistance to anticancer drugs. The most common ways to attain resistance in cancer cells include alteration in drug transport and metabolism, modification of drug target, elevated DNA damage response, or...
3.8K
Reinforcement
918
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
918
Corrosion of Reinforcement
577
The corrosion of steel reinforcement within concrete is a process influenced by the material's inherent properties and external factors. The high pH level of around 13, provided by calcium hydroxide present in concrete, initially protects the steel reinforcement by promoting the formation of a passive iron oxide layer on its surface.
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
577
Reinforcement Schedules
501
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
501
Reinforcements in Concrete
466
Reinforced concrete is a composite material used extensively in construction, combining the compressive strength of concrete with the tensile strength of steel. This synergy is essential as concrete, while excellent at resisting compression, is weak under tension. Steel bars, or rebars, are embedded in the concrete to handle these tensile forces. The choice of steel is strategic; it shares a similar coefficient of thermal expansion with concrete, which ensures uniformity in response to...
466
Machines
577
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. One example of a machine is the cutting plier, which is used to cut wires by applying forces to its handles. When equal and opposite forces are exerted on the handles of the cutting plier, they cause the cutting edges to come together and apply equal and opposite reaction forces on the wire, which are greater than the applied forces.
A free-body diagram of the...
A free-body diagram of the...
577

