Related Experiment Video
Updated: Nov 4, 2025

Evaluating the Effectiveness of Cancer Drug Sensitization In Vitro and In Vivo
Published on: February 6, 2015
The Unreasonable Effectiveness of Inverse Reinforcement Learning in Advancing Cancer Research
John Kalantari1,2, Heidi Nelson1,3, Nicholas Chia1,2,4
1Microbiome Program, Center for Individualized Medicine, Mayo Clinic, Rochester, MN, USA.
Abstract:
The "No Free Lunch" theorem states that for any algorithm, elevated performance over one class of problems is offset by its performance over another. Stated differently, no algorithm works for everything. Instead, designing effective algorithms often means exploiting prior knowledge of data relationships specific to a given problem. This "unreasonable efficacy" is especially desirable for complex and seemingly intractable problems in the natural sciences. One such area that is rife with the need for better algorithms is cancer biology-a field where relatively few insights are being generated from relatively large amounts of data. In part, this is due to the inability of mere statistics to reflect cancer as a genetic evolutionary process-one that involves cells actively mutating in order to navigate host barriers, outcompete neighboring cells, and expand spatially. Our work is built upon the central proposition that the Markov Decision Process (MDP) can better represent the process by which cancer arises and progresses. More specifically, by encoding a cancer cell's complex behavior as a MDP, we seek to model the series of genetic changes, or evolutionary trajectory, that leads to cancer as an optimal decision process. We posit that using an Inverse Reinforcement Learning (IRL) approach will enable us to reverse engineer an optimal policy and reward function based on a set of expert demonstrations extracted from the DNA of patient tumors. The inferred reward function and optimal policy can subsequently be used to extrapolate the evolutionary trajectory of any tumor. Here, we introduce a Bayesian nonparametric IRL model (PUR-IRL) where the number of reward functions is a priori unbounded in order to account for uncertainty in cancer data, i.e., the existence of latent trajectories and non-uniform sampling. We show that PUR-IRL is "unreasonably effective" in gaining interpretable and intuitive insights about cancer progression from high-dimensional genome data.
Related Concept Videos
Cancer Survival Analysis
Adaptive Mechanisms in Cancer Cells
Some of the advantages that cancer cells have on normal cells include - enhanced ability to divide without terminally differentiating, induce new blood vessel formation,...
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Treatment Resistant Cancers
Targeted Cancer Therapies
There are several types of targeted therapies against...
Cancer Vaccines
Cancer vaccines come in two categories: preventive (prophylactic) and treatment (active). Preventive vaccines, such as the Human Papillomavirus (HPV) vaccine, protect against viruses that cause certain...

