Related Experiment Video
Updated: May 14, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Confounding control in a nonexperimental study of STAR*D data: logistic regression balanced covariates better than
Alan R Ellis1, Stacie B Dusetzina, Richard A Hansen
1Cecil G. Sheps Center for Health Services Research, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599-7590, USA. are@unc.edu
Purpose:
Propensity scores (PSs), a powerful bias-reduction tool, can balance treatment groups on measured covariates in nonexperimental studies. We demonstrate the use of multiple PS estimation methods to optimize covariate balance.
Methods:
We used secondary data from 1292 adults with nonpsychotic major depressive disorder in the Sequenced Treatment Alternatives to Relieve Depression trial (2001-2004). After initial citalopram treatment failed, patient preference influenced assignment to medication augmentation (n = 565) or switch (n = 727). To reduce selection bias, we used boosted classification and regression trees (BCART) and logistic regression iteratively to identify two potentially optimal PSs. We assessed and compared covariate balance.
Results:
After iterative selection of interaction terms to minimize imbalance, logistic regression yielded better balance than BCART (average standardized absolute mean difference across 47 covariates: 0.03 vs. 0.08, matching; 0.02 vs. 0.05, weighting).
Conclusions:
Comparing multiple PS estimates is a pragmatic way to optimize balance. Logistic regression remains valuable for this purpose. Simulation studies are needed to compare PS models under varying conditions. Such studies should consider more flexible estimation methods, such as logistic models with automated selection of interactions or hybrid models using main effects logistic regression instead of a constant log-odds as the initial model for BCART.
More Related Videos
Related Concept Videos
Confounding in Epidemiological Studies
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Randomized Experiments
Simple randomization
Simple...

