Monotherapy cancer drug-blind response prediction is limited to intraclass generalization
William G Herbert1,2,3,4, Nicholas Chia5, Paul A Jensen6,7
1Graduate School of Biomedical Sciences, Mayo Clinic, Rochester, Minnesota, United States of America.
Abstract:
Monotherapy cancer drug response prediction (DRP) models predict the response of a cell line to a given drug. Analyzing these models' performance includes assessing their ability to predict the response of cell lines to new drugs, i.e., drugs that are not in the training set. Drug-blind prediction displays greatly diminished performance or outright failure across a wide range of model architectures and different large pharmacogenomic datasets. Drug-blind failure is hypothesized to be caused by the relatively limited set of drugs present in these datasets. The time and cost associated with further cell line experiments is significant, and it is impossible to predict beforehand how much data would be enough to overcome drug-blind failure. We must first define how current data contributes to drug-blind failure before attempting to remedy drug-blind failure with further data collection. In this work, we quantify the extent to which drug-blind generalizability relies on mechanistic overlap of drugs between training and testing splits. We first identify that the majority of mixed set DRP model performance can be attributed to drug overfitting, likely inhibiting generalization and preventing accurate analysis. Then, by specifically probing the drug-blind ability of models, we reveal the sources of generalizable drug features are confined to shared mechanisms of action and related pathways. Furthermore, we observed that, for certain mechanisms, we can significantly improve performance by limiting the training of models to a single mechanism compared to training on all drugs simultaneously. Across multiple different model architectures examined in this paper, we observe that drug-blind performance is a poor benchmark for DRP as it does not describe model behavior, it describes dataset behavior. Our investigation displays that these deep learning models trained on large, monotherapy cell line panels can more accurately describe mechanism of action of drugs rather than their advertised connection to broader cancer biology.
Insights
Drug-blind prediction models fail because they overfit to training drugs. Improved cancer drug response prediction requires focusing on shared drug mechanisms, not just broad cancer biology.
Area of Science:
- Computational biology
- Pharmacogenomics
- Machine learning in oncology
Background:
- Monotherapy cancer drug response prediction (DRP) models aim to predict cell line responses to drugs.
- Drug-blind prediction, assessing models on unseen drugs, reveals significantly diminished performance across various DRP models and datasets.
- This failure is often attributed to limited drug diversity in training datasets, making it difficult to generalize.
Purpose of the Study:
- To quantify the reliance of drug-blind generalizability on mechanistic overlap between training and testing drugs.
- To identify the sources of generalizable features in DRP models.
- To evaluate the effectiveness of training strategies focused on drug mechanisms.
Main Methods:
- Quantified drug-blind generalizability based on mechanistic overlap.
- Analyzed DRP model performance on training and testing splits with varying drug sets.
- Probed generalizable drug features by examining shared mechanisms of action and pathways.
- Compared training models on single mechanisms versus all drugs simultaneously.
Main Results:
- The majority of DRP model performance in mixed sets is due to drug overfitting, hindering generalization.
- Generalizable drug features are primarily linked to shared mechanisms of action and related pathways.
- Training models on single mechanisms can significantly improve performance for certain drug classes.
- Drug-blind performance is a poor benchmark for DRP, reflecting dataset characteristics more than model behavior.
Conclusions:
- Current DRP models primarily learn drug mechanisms of action rather than broader cancer biology.
- Drug-blind failure highlights the critical role of mechanistic understanding in developing generalizable DRP models.
- Future DRP model development should prioritize mechanistic diversity and targeted training strategies.
More Related Videos
Related Concept Videos
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Targeted Cancer Therapies
There are several types of targeted therapies against...
Treatment Resistant Cancers
Cancer Survival Analysis
Pharmacogenetics of Drug Targets: β₂-Adrenergic Receptors, Apo E, Thymidylate Synthase
Treatment Resistent Cancers


