Related Experiment Video
Updated: Jul 5, 2026

Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
Published on: September 17, 2019
Multiplicity in randomised trials II: subgroup and interim analyses
Kenneth F Schulz1, David A Grimes
1Family Health International, PO Box 13950, Research Triangle Park, NC 27709, USA. KSchulz@fhi.org
Abstract:
Subgroup analyses can pose serious multiplicity concerns. By testing enough subgroups, a false-positive result will probably emerge by chance alone. Investigators might undertake many analyses but only report the significant effects, distorting the medical literature. In general, we discourage subgroup analyses. However, if they are necessary, researchers should do statistical tests of interaction, rather than analyse every separate subgroup. Investigators cannot avoid interim analyses when data monitoring is indicated. However, repeatedly testing at every interim raises multiplicity concerns, and not accounting for multiplicity escalates the false-positive error. Statistical stopping methods must be used. The O'Brien-Fleming and Peto group sequential stopping methods are easily implemented and preserve the intended alpha level and power. Both adopt stringent criteria (low nominal p values) during the interim analyses. Implementing a trial under these stopping rules resembles a conventional trial, with the exception that it can be terminated early should a treatment prove greatly superior. Investigators and readers, however, need to grasp that the estimated treatment effects are prone to exaggeration, a random high, with early stopping.
Insights
Subgroup analyses and interim analyses can inflate false-positive rates. Use statistical interaction tests or group sequential methods like O'Brien-Fleming to manage multiplicity concerns and maintain study integrity.
Area of Science:
- Biostatistics
- Clinical Trial Design
- Medical Research Methodology
Background:
- Subgroup analyses and interim analyses in clinical trials present significant multiplicity concerns.
- Uncontrolled subgroup analyses can lead to spurious findings and distort the medical literature by selectively reporting significant results.
- Repeated interim analyses without proper statistical adjustment escalate the risk of false-positive errors.
Purpose of the Study:
- To address the multiplicity concerns associated with subgroup analyses and interim analyses in clinical trials.
- To recommend appropriate statistical methods for conducting necessary subgroup analyses and interim analyses.
- To highlight the importance of preserving statistical power and controlling the false-positive error rate.
Main Methods:
- Discouraging routine subgroup analyses, advocating for statistical tests of interaction when subgroup investigations are essential.
- Recommending the use of statistical stopping methods, such as O'Brien-Fleming and Peto group sequential methods, for interim analyses.
- Emphasizing the need to account for multiplicity to prevent escalating false-positive error rates.
Main Results:
- Subgroup analyses, if performed, should utilize interaction tests rather than analyzing each subgroup independently.
- Group sequential stopping methods effectively preserve the intended alpha level and statistical power during interim analyses.
- Early termination of trials due to apparent treatment superiority using these methods can lead to exaggerated treatment effect estimates.
Conclusions:
- Subgroup analyses should be approached with caution due to multiplicity concerns; interaction tests are preferred.
- Statistical stopping methods are crucial for managing multiplicity in interim analyses, ensuring trial integrity.
- While early stopping can be beneficial, researchers and readers must be aware of potential overestimation of treatment effects.
Related Concept Videos
Group Design
Introduction to Test of Independence
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes
Two-Way ANOVA
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...
Randomized Experiments
Simple randomization
Simple...

