Related Experiment Video
Updated: Dec 14, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Multiple-Imputation Variance Estimation in Studies With Missing or Misclassified Inclusion Criteria
Abstract:
In observational studies using routinely collected data, a variable with a high level of missingness or misclassification may determine whether an observation is included in the analysis. In settings where inclusion criteria are assessed after imputation, the popular multiple-imputation variance estimator proposed by Rubin ("Rubin's rules" (RR)) is biased due to incompatibility between imputation and analysis models. While alternative approaches exist, most analysts are not familiar with them. Using partially validated data from a human immunodeficiency virus cohort, we illustrate the calculation of an imputation variance estimator proposed by Robins and Wang (RW) in a scenario where the study exclusion criteria are based on a variable that must be imputed. In this motivating example, the corresponding imputation variance estimate for the log odds was 29% smaller using the RW estimator than using the RR estimator. We further compared these 2 variance estimators with a simulation study which showed that coverage probabilities of 95% confidence intervals based on the RR estimator were too high and became worse as more observations were imputed and more subjects were excluded from the analysis. The RW imputation variance estimator performed much better and should be employed when there is incompatibility between imputation and analysis models. We provide analysis code to aid future analysts in implementing this method.
Insights
When analyzing observational data, Rubin's rules for multiple imputation can be biased if exclusion criteria depend on imputed variables. The Robins-Wang estimator offers a more accurate alternative for variance estimation in such cases.
Area of Science:
- Biostatistics
- Epidemiology
- Data Science
Background:
- Routinely collected data in observational studies often contain missing or misclassified variables that influence analysis inclusion.
- Multiple imputation is common, but standard variance estimators like Rubin's rules (RR) can be biased when imputation and analysis models are incompatible, particularly with post-imputation exclusion criteria.
Purpose of the Study:
- To illustrate and evaluate the Robins-Wang (RW) imputation variance estimator as an alternative to Rubin's rules (RR) when study exclusion criteria rely on imputed variables.
- To compare the performance of RW and RR estimators using a human immunodeficiency virus (HIV) cohort dataset and a simulation study.
Main Methods:
- Applied the Robins-Wang (RW) imputation variance estimator to a partially validated HIV cohort dataset where exclusion criteria were based on an imputed variable.
- Conducted a simulation study to compare the bias and coverage probabilities of 95% confidence intervals generated by the RW and RR estimators under varying degrees of imputation and subject exclusion.
Main Results:
- The RW estimator yielded a 29% smaller imputation variance estimate for the log odds compared to the RR estimator in the HIV cohort example.
- Simulation results indicated that RR-based confidence intervals exhibited inflated coverage probabilities, which worsened with increased imputation and exclusion rates.
- The RW estimator demonstrated superior performance in maintaining accurate coverage probabilities.
Conclusions:
- The Robins-Wang (RW) imputation variance estimator should be preferred over Rubin's rules (RR) when imputation and analysis models are incompatible, especially when exclusion criteria depend on imputed data.
- The RW method provides a more reliable approach for variance estimation in complex observational studies using routinely collected data.
- Analysis code is provided to facilitate the adoption of the RW estimator by researchers.
Related Concept Videos
Confounding in Epidemiological Studies
Mechanistic Models: Compartment Models in Individual and Population Analysis
Comparing the Survival Analysis of Two or More Groups
Censoring Survival Data
Assumptions of Survival Analysis
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...

