A machine learning framework for supervised treatment response prediction from tumor transcriptomics: A large-scale
Lipika Ray Pal1,2, Edward Michael Gertz1,2, Nishanth Ulhas Nair1,2
1Cancer Data Science Laboratory (CDSL), Center for Cancer Research (CCR), National Cancer Institute (NCI), National Institutes of Health (NIH), Bethesda, MD, USA.
Abstract:
Precision oncology aims to guide treatment decisions using biomarkers. While DNA-based panels are increasingly applied, RNA transcriptomics remain underused due to limited datasets and the absence of robust models. We assembled the largest transcriptomic resource for drug response prediction to date, spanning 69 cohorts, 3,729 patients, nine cancer types, and six frontline therapies: anti-PD-1/PD-L1 immune-checkpoint inhibitors, trastuzumab, bevacizumab, BRAF inhibitors, paclitaxel, and FAC/FEC (Fluorouracil-Adriamycin-Cyclophosphamide/Fluorouracil-Epirubicin-Cyclophosphamide) chemotherapy. We developed EXPRESSO (EXpression-Profile-RESponSe-Optimizer), a supervised machine-learning framework that predicts treatment response from pre-treatment transcriptomes by integrating drug targets and context-specific biomarkers. EXPRESSO achieves ROC-AUCs of 0.64-0.73 and odds ratios of 2.4-4.6 across therapies, outperforming 20 published transcriptomic signatures. Robustness analysis reveals that predictive performance plateaued for some therapies with increasing training cohorts but continued to improve for others. These findings suggest inherent limits of supervised brute-force learning for certain treatments, but additional data and deeper mechanistic modeling may further enhance transcriptomics-based predictors.
Insights
This study introduces EXPRESSO, a machine-learning model using RNA data to predict cancer drug response. It shows promise for personalized medicine, outperforming existing methods and highlighting future research directions.
Area of Science:
- Oncology
- Genomics
- Bioinformatics
Background:
- Precision oncology leverages biomarkers for tailored cancer treatments.
- RNA transcriptomics are underutilized for drug response prediction due to data limitations and model scarcity.
Purpose of the Study:
- To develop a robust machine-learning framework for predicting patient response to various cancer therapies using pre-treatment transcriptomic data.
- To create the largest transcriptomic dataset for drug response prediction.
Main Methods:
- Assembled a large dataset (69 cohorts, 3,729 patients) across nine cancer types and six frontline therapies.
- Developed EXPRESSO (EXpression-Profile-RESponSe-Optimizer), a supervised machine-learning model integrating drug targets and biomarkers.
- Evaluated EXPRESSO's performance against 20 published transcriptomic signatures.
Main Results:
- EXPRESSO achieved significant predictive performance (ROC-AUCs 0.64–0.73, odds ratios 2.4–4.6) across multiple therapies.
- The model outperformed 20 existing transcriptomic signatures.
- Robustness analysis indicated performance plateaus for some therapies, suggesting potential limits of current supervised learning approaches.
Conclusions:
- Transcriptomic data, when analyzed with advanced machine learning like EXPRESSO, can effectively predict cancer treatment response.
- Further data and mechanistic modeling may enhance the predictive power of transcriptomic biomarkers for personalized oncology.
- EXPRESSO represents a significant advancement in utilizing RNA data for precision cancer therapy selection.
Related Concept Videos
Cancer Survival Analysis
Treatment Resistant Cancers
Tumor Immunotherapy


