Related Experiment Video
Updated: Jul 26, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
On the Q statistic with constant weights in meta-analysis of binary outcomes
Elena Kulinskaya1, David C Hoaglin2
1School of Computing Sciences, University of East Anglia, Norwich Research Park, NR4 7TJ, Norwich, UK. e.kulinskaya@uea.ac.uk.
Background:
Cochran's Q statistic is routinely used for testing heterogeneity in meta-analysis. Its expected value (under an incorrect null distribution) is part of several popular estimators of the between-study variance, [Formula: see text]. Those applications generally do not account for use of the studies' estimated variances in the inverse-variance weights that define Q (more explicitly, [Formula: see text]). Importantly, those weights make approximating the distribution of [Formula: see text] rather complicated.
Methods:
As an alternative, we are investigating a Q statistic, [Formula: see text], whose constant weights use only the studies' arm-level sample sizes. For log-odds-ratio (LOR), log-relative-risk (LRR), and risk difference (RD) as the measures of effect, we study, by simulation, approximations to distributions of [Formula: see text] and [Formula: see text], as the basis for tests of heterogeneity.
Results:
The results show that: for LOR and LRR, a two-moment gamma approximation to the distribution of [Formula: see text] works well for small sample sizes, and an approximation based on an algorithm of Farebrother is recommended for larger sample sizes. For RD, the Farebrother approximation works very well, even for small sample sizes. For [Formula: see text], the standard chi-square approximation provides levels that are much too low for LOR and LRR and too high for RD. The Kulinskaya et al. (Res Synth Methods 2:254-70, 2011) approximation for RD and the Kulinskaya and Dollinger (BMC Med Res Methodol 15:49, 2015) approximation for LOR work well for [Formula: see text] but have some convergence issues for very small sample sizes combined with small probabilities.
Conclusions:
The performance of the standard [Formula: see text] approximation is inadequate for all three binary effect measures. Instead, we recommend a test of heterogeneity based on [Formula: see text] and provide practical guidelines for choosing an appropriate test at the .05 level for all three effect measures.
Related Concept Videos
One-Way ANOVA: Unequal Sample Sizes
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Test for Homogeneity
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Friedman Two-way Analysis of Variance by Ranks

