Related Experiment Video
Updated: Dec 14, 2025

06:55
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
15.0K
Multiple-Imputation Variance Estimation in Studies With Missing or Misclassified Inclusion Criteria
American Journal of Epidemiology
|July 21, 2020
Summary
When analyzing observational data, Rubin's rules for multiple imputation can be biased if exclusion criteria depend on imputed variables. The Robins-Wang estimator offers a more accurate alternative for variance estimation in such cases.
Area of Science:
- Biostatistics
- Epidemiology
- Data Science
Background:
- Routinely collected data in observational studies often contain missing or misclassified variables that influence analysis inclusion.
- Multiple imputation is common, but standard variance estimators like Rubin's rules (RR) can be biased when imputation and analysis models are incompatible, particularly with post-imputation exclusion criteria.
Purpose of the Study:
- To illustrate and evaluate the Robins-Wang (RW) imputation variance estimator as an alternative to Rubin's rules (RR) when study exclusion criteria rely on imputed variables.
- To compare the performance of RW and RR estimators using a human immunodeficiency virus (HIV) cohort dataset and a simulation study.
Main Methods:
- Applied the Robins-Wang (RW) imputation variance estimator to a partially validated HIV cohort dataset where exclusion criteria were based on an imputed variable.
- Conducted a simulation study to compare the bias and coverage probabilities of 95% confidence intervals generated by the RW and RR estimators under varying degrees of imputation and subject exclusion.
Main Results:
- The RW estimator yielded a 29% smaller imputation variance estimate for the log odds compared to the RR estimator in the HIV cohort example.
- Simulation results indicated that RR-based confidence intervals exhibited inflated coverage probabilities, which worsened with increased imputation and exclusion rates.
- The RW estimator demonstrated superior performance in maintaining accurate coverage probabilities.
Conclusions:
- The Robins-Wang (RW) imputation variance estimator should be preferred over Rubin's rules (RR) when imputation and analysis models are incompatible, especially when exclusion criteria depend on imputed data.
- The RW method provides a more reliable approach for variance estimation in complex observational studies using routinely collected data.
- Analysis code is provided to facilitate the adoption of the RW estimator by researchers.
Related Concept Videos
Confounding in Epidemiological Studies
470
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
470
Mechanistic Models: Compartment Models in Individual and Population Analysis
182
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
182
Comparing the Survival Analysis of Two or More Groups
468
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
468
Censoring Survival Data
425
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
425
Assumptions of Survival Analysis
302
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
302
What are Estimates?
7.7K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
7.7K

