Related Experiment Video
Updated: Mar 6, 2026

Discovery of Driver Genes in Colorectal HT29-derived Cancer Stem-Like Tumorspheres
Published on: July 22, 2020
Systematic assessment of multi-gene predictors of pan-cancer cell line sensitivity to drugs exploiting gene
Linh Nguyen1, Cuong C Dang1, Pedro J Ballester1
1Cancer Research Center of Marseille, INSERM U1068, Marseille, France; Institut Paoli-Calmettes, Marseille, France; Aix-Marseille Université, Marseille, France; Cancer Research Center of Marseille UMR7258, Marseille, France.
Abstract:
Background: Selected gene mutations are routinely used to guide the selection of cancer drugs for a given patient tumour. Large pharmacogenomic data sets, such as those by Genomics of Drug Sensitivity in Cancer (GDSC) consortium, were introduced to discover more of these single-gene markers of drug sensitivity. Very recently, machine learning regression has been used to investigate how well cancer cell line sensitivity to drugs is predicted depending on the type of molecular profile. The latter has revealed that gene expression data is the most predictive profile in the pan-cancer setting. However, no study to date has exploited GDSC data to systematically compare the performance of machine learning models based on multi-gene expression data against that of widely-used single-gene markers based on genomics data. Methods: Here we present this systematic comparison using Random Forest (RF) classifiers exploiting the expression levels of 13,321 genes and an average of 501 tested cell lines per drug. To account for time-dependent batch effects in IC 50 measurements, we employ independent test sets generated with more recent GDSC data than that used to train the predictors and show that this is a more realistic validation than standard k-fold cross-validation. Results and Discussion: Across 127 GDSC drugs, our results show that the single-gene markers unveiled by the MANOVA analysis tend to achieve higher precision than these RF-based multi-gene models, at the cost of generally having a poor recall (i.e. correctly detecting only a small part of the cell lines sensitive to the drug). Regarding overall classification performance, about two thirds of the drugs are better predicted by the multi-gene RF classifiers. Among the drugs with the most predictive of these models, we found pyrimethamine, sunitinib and 17-AAG. Conclusions: Thanks to this unbiased validation, we now know that this type of models can predict in vitro tumour response to some of these drugs. These models can thus be further investigated on in vivo tumour models. R code to facilitate the construction of alternative machine learning models and their validation in the presented benchmark is available at http://ballester.marseille.inserm.fr/gdsc.transcriptomicDatav2.tar.gz.
Insights
Machine learning models using multi-gene expression data predict cancer drug sensitivity better than single-gene markers. This study compared Random Forest models against traditional methods, finding multi-gene approaches more effective for predicting tumor response.
Area of Science:
- Genomics
- Pharmacogenomics
- Machine Learning
Background:
- Gene mutations guide cancer drug selection.
- Pharmacogenomic datasets like GDSC aim to find single-gene drug sensitivity markers.
- Machine learning shows gene expression is a predictive profile for cancer drug sensitivity.
Purpose of the Study:
- Systematically compare machine learning models using multi-gene expression data against single-gene markers from genomics data.
- Evaluate the performance of Random Forest classifiers on large-scale pharmacogenomic data.
- Assess the predictive power of multi-gene expression models for cancer cell line drug sensitivity.
Main Methods:
- Utilized Random Forest (RF) classifiers with expression levels of 13,321 genes.
- Employed independent test sets from recent GDSC data for realistic validation, accounting for batch effects.
- Compared RF models against single-gene markers identified via MANOVA analysis across 127 GDSC drugs.
Main Results:
- Single-gene markers showed higher precision but lower recall compared to RF multi-gene models.
- Multi-gene RF classifiers achieved better overall classification performance for approximately two-thirds of the drugs.
- Pyrimethamine, sunitinib, and 17-AAG were among the drugs best predicted by multi-gene RF models.
Conclusions:
- Unbiased validation confirms multi-gene expression models can predict in vitro tumor response to certain drugs.
- These machine learning models warrant further investigation in in vivo tumor models.
- R code is available for constructing and validating alternative machine learning models.

