Related Experiment Video
Updated: Aug 5, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Evaluating large-scale propensity score adjustment when sample size is small
Fleur Vereijken1, Jenna M Reps1,2, Marc A Suchard3
1Department of Medical Informatics, Erasmus University Medical Center, Rotterdam, 3015 GD, The Netherlands.
Large-scale propensity score (LSPS) models effectively reduce bias in small observational studies. While standard diagnostics may fail, LSPS models support reliable confounding control in federated analyses.
Area of Science:
- Epidemiology
- Biostatistics
- Health Data Science
Background:
- Large-scale propensity score (LSPS) models are vital for controlling confounding in observational research.
- Concerns exist regarding the reliability of LSPS in small-sample settings due to potential issues with accurate propensity score estimation and covariate balance.
- Small datasets are often excluded, losing valuable information.
Purpose of the Study:
- To evaluate the performance of LSPS models in small-sample settings.
- To assess the balancing and bias reduction capabilities of LSPS models under varying sample sizes.
- To compare local LSPS models within data subsets to a global LSPS model trained on the full dataset.
Main Methods:
- LSPS models were evaluated using six target-comparator pairs across partitioned real-world data sources.
- Performance was assessed based on covariate balance, treatment effect bias, and precision after empirical calibration.
- Methods included real negative controls and synthetic positive controls, emulating a distributed data network analysis.
Main Results:
- LSPS models significantly reduced bias and improved precision with empirical calibration.
- Propensity score (PS) matching demonstrated greater resilience to small sample sizes than PS stratification.
- Standard balance diagnostics were unreliable below 20,000 participants, unlike a chance imbalance diagnostic.
Conclusions:
- LSPS models offer reliable bias reduction in small-sample settings, suitable for federated analyses.
- Standard balance diagnostics can be misleading in small samples; alternatives like significance checking are recommended.
- When LSPS models inadequately reduce bias, supplementary adjustment strategies are necessary.
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Sample Proportion and Population Proportion
One-Way ANOVA: Unequal Sample Sizes
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Distributions to Estimate Population Parameter
