Related Experiment Video
Updated: Apr 17, 2026

12:18
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
8.2K
A regression tree approach to identifying subgroups with differential treatment effects
Wei-Yin Loh1, Xu He, Michael Man
1Department of Statistics, University of Wisconsin, Madison, WI, 53706, U.S.A.
Statistics in Medicine
|February 7, 2015
Summary
New regression tree algorithms identify patient subgroups benefiting from cancer treatments. These methods reduce bias in clinical trials, improving treatment efficacy discovery for regulatory approval.
Area of Science:
- Biostatistics
- Clinical Trial Design
- Machine Learning in Medicine
Background:
- Discovering effective treatments for complex diseases like cancer is challenging.
- Identifying patient subgroups with enhanced treatment response is crucial for regulatory approval.
Purpose of the Study:
- Introduce novel regression tree algorithms for subgroup identification.
- Address limitations of existing methods in handling complex clinical trial data.
Main Methods:
- Extend the generalized unbiased interaction detection and estimation (GUIDE) approach.
- Incorporate treatment as a linear predictor, chi-squared tests, and Poisson regression for proportional hazards modeling.
- Handle censored data and missing predictor variables.
Main Results:
- Developed three new regression tree algorithms practically free of selection bias.
- Algorithms are applicable to randomized trials with multiple treatments and complex data.
- Generated importance scores and confidence intervals for treatment effects.
Conclusions:
- The new algorithms offer a robust approach for identifying patient subgroups with differential treatment effects.
- These methods enhance the analysis of clinical trial data, aiding in personalized medicine strategies.
- The techniques are validated using both real and simulated datasets.
Related Concept Videos
Survival Tree
504
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
504
Comparing the Survival Analysis of Two or More Groups
727
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
727
Regression Toward the Mean
7.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.3K
Regression Analysis
9.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
9.1K
Cochran's Q Test
1.2K
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
1.2K
Quantifying and Rejecting Outliers: The Grubbs Test
5.0K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
5.0K

