Related Experiment Video
Updated: Apr 17, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
A regression tree approach to identifying subgroups with differential treatment effects
Wei-Yin Loh1, Xu He, Michael Man
1Department of Statistics, University of Wisconsin, Madison, WI, 53706, U.S.A.
Abstract:
In the fight against hard-to-treat diseases such as cancer, it is often difficult to discover new treatments that benefit all subjects. For regulatory agency approval, it is more practical to identify subgroups of subjects for whom the treatment has an enhanced effect. Regression trees are natural for this task because they partition the data space. We briefly review existing regression tree algorithms. Then, we introduce three new ones that are practically free of selection bias and are applicable to data from randomized trials with two or more treatments, censored response variables, and missing values in the predictor variables. The algorithms extend the generalized unbiased interaction detection and estimation (GUIDE) approach by using three key ideas: (i) treatment as a linear predictor, (ii) chi-squared tests to detect residual patterns and lack of fit, and (iii) proportional hazards modeling via Poisson regression. Importance scores with thresholds for identifying influential variables are obtained as by-products. A bootstrap technique is used to construct confidence intervals for the treatment effects in each node. The methods are compared using real and simulated data.
Insights
New regression tree algorithms identify patient subgroups benefiting from cancer treatments. These methods reduce bias in clinical trials, improving treatment efficacy discovery for regulatory approval.
Area of Science:
- Biostatistics
- Clinical Trial Design
- Machine Learning in Medicine
Background:
- Discovering effective treatments for complex diseases like cancer is challenging.
- Identifying patient subgroups with enhanced treatment response is crucial for regulatory approval.
Purpose of the Study:
- Introduce novel regression tree algorithms for subgroup identification.
- Address limitations of existing methods in handling complex clinical trial data.
Main Methods:
- Extend the generalized unbiased interaction detection and estimation (GUIDE) approach.
- Incorporate treatment as a linear predictor, chi-squared tests, and Poisson regression for proportional hazards modeling.
- Handle censored data and missing predictor variables.
Main Results:
- Developed three new regression tree algorithms practically free of selection bias.
- Algorithms are applicable to randomized trials with multiple treatments and complex data.
- Generated importance scores and confidence intervals for treatment effects.
Conclusions:
- The new algorithms offer a robust approach for identifying patient subgroups with differential treatment effects.
- These methods enhance the analysis of clinical trial data, aiding in personalized medicine strategies.
- The techniques are validated using both real and simulated datasets.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Comparing the Survival Analysis of Two or More Groups
Regression Toward the Mean
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Cochran's Q Test
Quantifying and Rejecting Outliers: The Grubbs Test

