Related Experiment Video
Updated: Aug 5, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Evaluating large-scale propensity score adjustment when sample size is small
Fleur Vereijken1, Jenna M Reps1,2, Marc A Suchard3
1Department of Medical Informatics, Erasmus University Medical Center, Rotterdam, 3015 GD, The Netherlands.
Background:
Large-scale propensity score (LSPS) models are increasingly used to control confounding in observational studies, but their reliability in small-sample settings is unclear. Small samples can limit a study's ability to estimate propensity scores accurately and achieve adequate covariate balance, raising concerns about insufficient confounding adjustment. Resulting in small datasets being excluded, despite potentially containing valid information.
Methods:
We evaluated the balancing and bias reduction performance of LSPS models using six target-comparator pairs. To emulate a distributed data network analysis setting, we partitioned two large real-world data sources into smaller subsets. Within each subset, LSPS models were fit locally and evaluated on covariate balance, treatment effect bias and precision after empirical calibration, using real negative controls and synthetic positive controls. Performance was compared to a global LSPS model trained on the full dataset.
Results:
Across most target-comparator pairs, LSPS models substantially reduced bias and, when applying empirical calibration, improved precision. PS matching was more resilient to small sample sizes than PS stratification, achieving adequate balance even at n = 500, while stratification sometimes failed at n = 4000. A slight bias increase was observed at the smallest sample sizes, though not universally. Standard balance diagnostics consistently failed below n = 20 000, while a recently proposed diagnostic accounting for chance imbalance did not.
Conclusions:
LSPS models generally provide reliable bias reduction in small-sample settings, supporting their use in federated analyses. However, standard balance diagnostics may be misleading in small samples, and alternatives should be considered, such as significance checking. When LSPS fails to reduce bias adequately, additional adjustment strategies are required.
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Sample Proportion and Population Proportion
One-Way ANOVA: Unequal Sample Sizes
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Distributions to Estimate Population Parameter
