Related Experiment Video
Updated: Sep 19, 2025

Author Spotlight: Evaluating the Adjuvant Efficacy and Safety of Angong Niuhuang Pill in Viral Encephalitis Treatment
Published on: April 19, 2024
The performance of small sample correction methods for controlling type I error when analyzing parallel cluster
K Hemming1, J Thompson1, C Kristunas2
1Applied Health Research, School of Health Sciences, College of Medicine and Health, University of Birmingham, Birmingham, UK.
Objectives:
Most cluster randomized trials (CRTs) include fewer than 50 clusters, yet the assumption of a "large sample" is relied upon to derive the sampling distributions of treatment effects. We review the current simulation study literature pertaining to small sample corrections for common analytical approaches for parallel CRTs.
Study Design And Setting:
We searched Ovid Medline and Web of Science up to August 30 2024 for simulation studies evaluating the performance of small sample corrections. We include binary and continuous outcomes analyzed using generalized linear mixed models, generalized estimating equations, or cluster-level approaches. Full-text screening and data abstraction was performed independently in duplicate.
Results:
Fourteen studies evaluated binary outcomes and 6 studies continuous outcomes. The number of clusters ranged from 4 to 200; the median smallest intracluster correlation coefficient was 0.001 [range 0.000-0.200]; largest intracluster correlation coefficient was 0.10 [range 0.05-0.70]; lowest prevalence was 0.25 [range 0.05-0.50]; and median coefficient of variation of cluster sizes 1.00 [range: 0.80-1.50]. For continuous outcomes, a cluster-level analysis (either unweighted or inverse-variance weighted) with a t-distribution (with between-within degree of freedom); a linear mixed model with a Satterthwaite correction; or a generalized estimating equation with the Fay and Graubard correction mostly preserve nominal type I error with as few as six clusters (although up to 40 clusters in some settings). Other approaches work less favorably (eg, Kenward-Roger is conservative even with 30 clusters). For binary outcomes, an unweighted or inverse-variance weighted cluster-level analysis can achieve nominal type I error (but can be anticonservative with small cluster sizes or low prevalence); as can a generalized linear mixed model with a between-within correction with as few as 10 clusters (but sometimes conservative with up to 30 clusters). Other corrections such as the Kenward-Roger or Satterthwaite correction are more conservative. For generalized estimating equations, the Mancl and DeRouen correction mostly seems to preserve nominal errors but can be anticonservative.
Conclusion:
The literature on the performance of small sample corrections for parallel CRTs is complex. While the available corrections can maintain type I error with a very small number of clusters, more than 40 clusters are required to guarantee nominal type I error across all settings.
Related Concept Videos
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Randomized Experiments
Simple randomization
Simple...
Bonferroni Test
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...

