RAINFOREST: a random forest approach to predict treatment benefit in data from (failed) clinical drug trials
Joske Ubels1,2,3,4, Tilman Schaefers1,4, Cornelis Punt5
1Center for Molecular Medicine, UMC Utrecht, Utrecht, The Netherlands.
Motivation:
When phase III clinical drug trials fail their endpoint, enormous resources are wasted. Moreover, even if a clinical trial demonstrates a significant benefit, the observed effects are often small and may not outweigh the side effects of the drug. Therefore, there is a great clinical need for methods to identify genetic markers that can identify subgroups of patients which are likely to benefit from treatment as this may (i) rescue failed clinical trials and/or (ii) identify subgroups of patients which benefit more than the population as a whole. When single genetic biomarkers cannot be found, machine learning approaches that find multivariate signatures are required. For single nucleotide polymorphism (SNP) profiles, this is extremely challenging owing to the high dimensionality of the data. Here, we introduce RAINFOREST (tReAtment benefIt prediction using raNdom FOREST), which can predict treatment benefit from patient SNP profiles obtained in a clinical trial setting.
Results:
We demonstrate the performance of RAINFOREST on the CAIRO2 dataset, a phase III clinical trial which tested the addition of cetuximab treatment for metastatic colorectal cancer and concluded there was no benefit. However, we find that RAINFOREST is able to identify a subgroup comprising 27.7% of the patients that do benefit, with a hazard ratio of 0.69 (P = 0.04) in favor of cetuximab. The method is not specific to colorectal cancer and could aid in reanalysis of clinical trial data and provide a more personalized approach to cancer treatment, also when there is no clear link between a single variant and treatment benefit.
Availability And Implementation:
The R code used to produce the results in this paper can be found at github.com/jubels/RAINFOREST. A more configurable, user-friendly Python implementation of RAINFOREST is also provided. Due to restrictions based on privacy regulations and informed consent of participants, phenotype and genotype data of the CAIRO2 trial cannot be made freely available in a public repository. Data from this study can be obtained upon request. Requests should be directed toward Prof. Dr. H.J. Guchelaar (h.j.guchelaar@lumc.nl).
Supplementary Information:
Supplementary data are available at Bioinformatics online.
Insights
RAINFOREST identifies patient subgroups benefiting from cancer drugs, rescuing failed trials and personalizing treatment. This machine learning approach analyzes single nucleotide polymorphism (SNP) profiles for improved clinical trial outcomes.
Area of Science:
- Genomics
- Bioinformatics
- Machine Learning
Background:
- Phase III clinical trials often fail to demonstrate drug efficacy, wasting resources.
- Even successful trials may show small benefits that don't outweigh side effects.
- Identifying patient subgroups who benefit most is crucial for personalized medicine and trial rescue.
Purpose of the Study:
- To introduce RAINFOREST (tReAtment benefIt prediction using raNdom FOREST), a machine learning method for predicting treatment benefit from patient single nucleotide polymorphism (SNP) profiles.
- To address the challenge of high-dimensional SNP data in identifying multivariate genetic signatures for treatment response.
- To rescue failed clinical trials and identify patient subgroups with superior treatment benefit.
Main Methods:
- RAINFOREST utilizes a machine learning approach, specifically random forest, to analyze high-dimensional SNP data.
- The method is designed to identify multivariate signatures predictive of treatment benefit.
- It was applied to the CAIRO2 dataset, a phase III clinical trial for metastatic colorectal cancer.
Main Results:
- RAINFOREST identified a subgroup of 27.7% of patients who benefited from cetuximab treatment in the CAIRO2 trial, despite the trial's overall negative conclusion.
- This subgroup showed a significant hazard ratio of 0.69 (P=0.04) in favor of cetuximab.
- The method demonstrated its ability to find treatment benefits even without a clear link between a single genetic variant and efficacy.
Conclusions:
- RAINFOREST can identify patient subgroups that benefit from treatments, even when overall trial results are negative.
- The method offers a powerful tool for re-analyzing clinical trial data and advancing personalized cancer treatment.
- RAINFOREST is applicable beyond colorectal cancer and aids in situations where single biomarkers are insufficient.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Related Concept Videos
Clinical Trials
There are four phases in a clinical trial. A phase one...
Clinical Trials: Overview
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Hazard Ratio
For example, in a clinical trial...
Regression Toward the Mean
Analysis of Population Pharmacokinetic Data
