Related Experiment Video
Updated: Mar 9, 2026

14:14
The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
Published on: May 13, 2022
6.4K
Common reporting errors in subgroup analysis: a comparison of interaction and stratified regression models.
1School of Social and Behavioral Sciences, Nanjing University, Nanjing, Jiangsu, 210023, China.
BMC Medical Research Methodology
|March 7, 2026
Summary
Researchers should avoid combining stratified and interaction regression for subgroup analysis. Interaction regression generally offers higher statistical power, but stratified regression better controls Type I error in specific scenarios.
Area of Science:
- Epidemiology
- Public Health
- Medical Research
Background:
- Reporting treatment effect estimates from stratified regressions alongside significance test results from interaction regressions is common practice.
- However, combining these methods raises methodological concerns due to differing theoretical foundations and statistical properties.
- This study evaluates their effectiveness in estimating heterogeneous treatment effects.
Purpose of the Study:
- To compare the performance of interaction regression and stratified regression in estimating heterogeneous treatment effects.
- To identify scenarios where each method excels or struggles.
- To provide guidance on appropriate reporting strategies for subgroup analyses.
Main Methods:
- Conducted Monte Carlo simulations varying sample size, subgroup proportions, covariate-outcome heterogeneity, and covariate correlation.
- Evaluated performance using mean squared error (MSE), empirical coverage rate (ECR), and empirical statistical power (ESP).
- Applied both methods to real-world data from the International Social Survey Programme (ISSP) for illustrative purposes.
Main Results:
- Stratified regression effectively controlled Type I error with large sample sizes and balanced groups.
- Interaction regression struggled with Type I error control when baseline characteristics differed and covariates were weakly correlated.
- Interaction regression generally outperformed stratified regression in other scenarios, offering higher statistical power with acceptable Type I error rates.
Conclusions:
- Interaction regression and stratified regression have distinct strengths and limitations for estimating heterogeneous treatment effects.
- Researchers should avoid hybrid reporting strategies combining estimates from one method with significance tests from another.
- Differences in Type I error control underscore the need for careful method selection and reporting in subgroup analyses.
Keywords:
Heterogeneous treatment effectsInteraction regressionMonte Carlo simulationsStratified regressionSubgroup analysesType I error controlMore Related Videos
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
677
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
677
Two-Way ANOVA
3.5K
The two-way ANOVA is an extension of the one-way ANOVA. It is a statistical test performed on three or more samples categorized by two factors - a row factor and a column factor. Ronald Fischer mentioned it in 1925 in his book 'Statistical Methods for Researchers.'
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
3.5K
Errors In Hypothesis Tests
6.1K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
6.1K
Confounding in Epidemiological Studies
919
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
919
Regression Analysis
8.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.7K
Friedman Two-way Analysis of Variance by Ranks
530
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
530

