Related Experiment Video
Updated: Sep 19, 2026

Dynamic Lung Tumor Tracking for Stereotactic Ablative Body Radiation Therapy
Published on: June 7, 2015
A Modular Multi-Agent Reinforcement Learning Framework Guided by LLMs: Improving Quality and Efficiency of Treatment
Zipai Wang1, Hao Guo1, Yang Lei1
1Department of Radiation Oncology, Icahn School of Medicine at Mount Sinai, New York, USA.
Purpose:
Knowledge-based planning (KBP) frequently requires manual refinement to satisfy institution-specific clinical constraints, while existing automated approaches based on large language models (LLMs) or reinforcement learning (RL) each face distinct limitations in long-term optimization awareness and scalability. We propose a hybrid LLM-guided modular RL framework that integrates the strengths of both approaches for automated lung IMRT treatment planning.
Methods:
The framework employs a two-level hierarchical architecture in which task-specific RL agents execute dose modifications under the coordination of an LLM supervisory layer comprising Planner, Evaluator, and Supervisor agents. Modular parallel and serial OAR agents are trained using SARSA with linear function approximation, with category-level weight sharing enabling cross-OAR generalization from a minimal training set. Upon achieving full constraint compliance, the system optionally enters Lung Sparing Mode, in which the Lung RL agent continues to reduce lung dose through single-step LLM-supervised iterations. The two hybrid configurations utilizing the proposed framework were evaluated on 62 retrospective LA-NSCLC cases against an institutional KBP model and an LLM-only system using institutional dose-volume constraints.
Results:
All three LLM-guided systems achieved a 98% clinical goal achievement rate, compared with 74% for KBP alone. Among the 16 cases requiring post-KBP refinement, the hybrid framework reached the clinical goal with 75% fewer LLM supervisory calls than the LLM-only system (1.7 ± 0.7 vs 6.7 ± 5.7) and in less time (20.6 ± 15.7 vs 33.4 ± 32.1 min), with the worst case reduced from 109.8 to 58.5 min. Enabling Lung Sparing Mode further reduced lung V20 (21.3 ± 8.4% vs 25.2 ± 10.1%, p < 0.001) and mean lung dose (14.2 ± 4.4 vs 15.0 ± 4.8 Gy, p < 0.001) relative to Normal Mode, at the cost of small increases in spinal cord and esophagus maximum dose that remained well within constraints. The resulting plans achieved lower lung dose than the clinically delivered plans for the same patients (V20 21.3% vs 24.6%; mean 14.2 vs 14.9 Gy).
Conclusions:
Integrating modular RL optimization with LLM-based supervision improves both lung sparing and planning efficiency relative to LLM-only automated planning, while enabling scalable deployment through cross-OAR weight sharing with reduced training requirements.
