Related Experiment Video
Updated: Jul 5, 2026

06:52
Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
Published on: September 17, 2019
Multiplicity in randomised trials II: subgroup and interim analyses.
Kenneth F Schulz1, David A Grimes
1Family Health International, PO Box 13950, Research Triangle Park, NC 27709, USA. KSchulz@fhi.org
Lancet (London, England)
|May 12, 2005
Summary
Subgroup analyses and interim analyses can inflate false-positive rates. Use statistical interaction tests or group sequential methods like O'Brien-Fleming to manage multiplicity concerns and maintain study integrity.
Area of Science:
- Biostatistics
- Clinical Trial Design
- Medical Research Methodology
Background:
- Subgroup analyses and interim analyses in clinical trials present significant multiplicity concerns.
- Uncontrolled subgroup analyses can lead to spurious findings and distort the medical literature by selectively reporting significant results.
- Repeated interim analyses without proper statistical adjustment escalate the risk of false-positive errors.
Purpose of the Study:
- To address the multiplicity concerns associated with subgroup analyses and interim analyses in clinical trials.
- To recommend appropriate statistical methods for conducting necessary subgroup analyses and interim analyses.
- To highlight the importance of preserving statistical power and controlling the false-positive error rate.
Main Methods:
- Discouraging routine subgroup analyses, advocating for statistical tests of interaction when subgroup investigations are essential.
- Recommending the use of statistical stopping methods, such as O'Brien-Fleming and Peto group sequential methods, for interim analyses.
- Emphasizing the need to account for multiplicity to prevent escalating false-positive error rates.
Main Results:
- Subgroup analyses, if performed, should utilize interaction tests rather than analyzing each subgroup independently.
- Group sequential stopping methods effectively preserve the intended alpha level and statistical power during interim analyses.
- Early termination of trials due to apparent treatment superiority using these methods can lead to exaggerated treatment effect estimates.
Conclusions:
- Subgroup analyses should be approached with caution due to multiplicity concerns; interaction tests are preferred.
- Statistical stopping methods are crucial for managing multiplicity in interim analyses, ensuring trial integrity.
- While early stopping can be beneficial, researchers and readers must be aware of potential overestimation of treatment effects.
Related Concept Videos
Group Design
The most basic experimental design involves two groups: the experimental group and the control group. The two groups are designed to be the same except for one difference— experimental manipulation. The experimental group gets the experimental manipulation—that is, the treatment or variable being tested—and the control group does not. Since experimental manipulation is the only difference between the experimental and control groups, we can be sure that any differences between the two are due to...
Introduction to Test of Independence
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
One-Way ANOVA: Equal Sample Sizes
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
Two-Way ANOVA
The two-way ANOVA is an extension of the one-way ANOVA. It is a statistical test performed on three or more samples categorized by two factors - a row factor and a column factor. Ronald Fischer mentioned it in 1925 in his book 'Statistical Methods for Researchers.'
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...
Randomized Experiments
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...

