Related Experiment Video
Updated: Jan 22, 2026

Methods of Soil Resampling to Monitor Changes in the Chemical Concentrations of Forest Soils
Published on: November 25, 2016
Predicting drug activity against cancer cells by random forest models based on minimal genomic information and
Alex P Lind1, Peter C Anderson1
1Physical Sciences Division, University of Washington Bothell, Bothell, Washington, United States of America.
Abstract:
A key goal of precision medicine is predicting the best drug therapy for a specific patient from genomic information. In oncology, cancers that appear similar pathologically can vary greatly in how they respond to the same drug. Fortunately, data from high-throughput screening programs often reveal important relationships between genomic variability of cancer cells and their response to drugs. Nevertheless, many current computational methods to predict compound activity against cancer cells require large quantities of genomic, epigenomic, and additional cellular data to develop and to apply. Here we integrate recent screening data and machine learning to train classification models that predict the activity/inactivity of compounds against cancer cells based on the mutational status of only 145 oncogenes and a set of compound structural descriptors. Using IC50 values of 1 μM as activity cutoffs, our predictive models have sensitivities of 87%, specificities of 87%, and yield an area under the receiver operating characteristic curve equal to 0.94. We also develop regression models to predict log(IC50) values of compounds for cancer cells; the models achieve a Pearson correlation coefficient of 0.86 for cross-validation and up to 0.65-0.73 against blind test sets. Predictive performance remains strong when as few as 50 oncogenes are included. Finally, even when 40% of experimental IC50 values are missing from screening data, they can be imputed with sufficient reliability that classification accuracy is not diminished. The presented models are fast to generate and may serve as easily implemented screening tools for personalized oncology medicine, drug repurposing, and drug discovery.
Insights
This study developed machine learning models to predict cancer drug response using genomic data from 145 oncogenes. These models accurately predict compound activity, aiding personalized oncology and drug discovery.
Area of Science:
- Computational biology
- Genomics
- Pharmacology
Background:
- Precision medicine aims to tailor drug therapies using genomic information.
- Cancer drug response varies significantly despite similar pathology.
- Existing computational methods often require extensive genomic and cellular data.
Purpose of the Study:
- To develop machine learning models for predicting anti-cancer compound activity.
- To utilize genomic data and compound structures for accurate predictions.
- To create efficient screening tools for personalized oncology.
Main Methods:
- Integrated screening data and machine learning to train classification and regression models.
- Used mutational status of 145 oncogenes and compound structural descriptors.
- Employed IC50 values as activity cutoffs and log(IC50) for regression.
Main Results:
- Classification models achieved 87% sensitivity and 87% specificity (AUC 0.94).
- Regression models yielded a Pearson correlation coefficient of 0.86 (cross-validation) and 0.65-0.73 (blind tests).
- Performance remained robust with as few as 50 oncogenes and with imputed missing data.
Conclusions:
- The developed models are fast and can be easily implemented.
- These models can serve as screening tools for personalized oncology, drug repurposing, and discovery.
- Accurate prediction of drug response from limited genomic data is feasible.
More Related Videos
Related Concept Videos
Factors Affecting Drug Biotransformation: Physicochemical and Chemical Properties of Drugs
The drug's acidity or basicity is essential in...
Physical and Chemical Properties of Matter
Urine: Physical and Chemical Properties
The concentration of these solutes varies, with urea being the most abundant nitrogenous waste product. Other solutes include sodium, chloride, potassium, phosphate,...
Genomics
Properties of Enantiomers and Optical Activity
Thermodynamics: Chemical Potential and Activity
The thermodynamic equilibrium constant is more accurately defined in terms of activity rather than concentration.

