Related Experiment Video
Updated: Aug 6, 2026

Modeling Chemotherapy Resistant Leukemia In Vitro
Published on: February 9, 2016
Reinforcement learning for chemotherapy scheduling in a stochastic tumor evolution model
1University of Southern California, Department of Aerospace & Mechanical Engineering, Los Angeles, California 90089-1191, USA.
Abstract:
We present a Q-learning framework for optimizing chemotherapy dosing schedules in a stochastic finite-cell model of tumor evolution under drug-induced selection. The tumor consists of three competing subpopulations: a chemosensitive lineage, S, and two single-drug-resistant lineages, R_{1} and R_{2}, each resistant to one of two cytotoxic agents, C_{1} and C_{2}. Drug administration is formulated as a discrete action space in a finite-state Markov process, where the state is defined by the composition (S,R_{1},R_{2}) constrained by a fixed population size N. Tumor volume evolves according to a separate growth equation, with expansion rate proportional to the difference between the population-averaged fitness and a fixed microenvironmental baseline. Using Q-learning, we derive optimal dosing policies that balance therapeutic pressure with the evolutionary dynamics of resistance. The reward function is engineered to promote long-term coexistence among subpopulations, thereby delaying fixation of resistance by penalizing population imbalance. We analyze the structure of the optimal policies to (i) infer dominant evolutionary trajectories under treatment, (ii) quantify robustness to partial observability of both initial conditions and state updates, and (iii) construct simplified, symmetry-informed heuristics that approximate the full learned policy. Our results highlight the potential of model-free adaptive control strategies to steer tumor evolution away from drug resistance in the presence of biological stochasticity and information constraints.
