Related Experiment Video
Updated: Sep 10, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Comparing Multiple Imputation Methods to Address Missing Patient Demographics in Immunization Information Systems:
Sara Brown1, Ousswa Kudia1, Kaye Kleine1
1Scientific Services - Analytics, Scientific Technologies Corporation (United States), 411 S 1st St, Phoenix, AZ, 85004, United States, 1 480-745-8500.
Multiple imputation methods like MICE and miceforest effectively manage missing race/ethnicity data in immunization surveillance, preserving demographics and improving accuracy for public health interventions. Miceforest offers better computational efficiency for large datasets.
Area of Science:
- Public Health Surveillance
- Biostatistics
- Health Informatics
Background:
- Immunization Information Systems (IIS) and surveillance data are crucial for public health but often suffer from missing data, potentially biasing vaccine coverage assessments and hindering efforts to address health disparities.
- Accurate assessment of vaccine coverage is essential for effective public health programming and interventions, especially when aiming to reduce inequities.
Purpose of the Study:
- To evaluate the performance of three multiple imputation methods—MICE, Iterative-Imputer, and miceforest—in handling missing race and ethnicity data within large-scale public health surveillance datasets.
- To compare these imputation methods based on their ability to preserve demographic distributions, computational efficiency, and their impact on assessing the association between race/ethnicity and flu vaccination status.
Main Methods:
- A retrospective cohort study analyzed 2021-2022 flu vaccination and demographic data from the West Virginia Immunization Information System (N=2,302,036), with significant missingness in race (15%) and ethnicity (34%).
- Three multiple imputation techniques (MICE, Iterative-Imputer, miceforest) were applied to generate 15 imputed datasets each.
- Performance was assessed by comparing demographic distribution preservation, computational efficiency, and spatial clustering patterns using G-statistics and likelihood ratio statistics.
Main Results:
- All imputation methods showed significant spatial clustering for race imputation. MICE and miceforest demonstrated superior preservation of demographic proportional distributions compared to Iterative-Imputer.
- Computational efficiency varied significantly: MICE took 14 hours, Iterative-Imputer 2 minutes, and miceforest 10 minutes for 15 imputations.
- Post-imputation analysis revealed reductions in stratified flu vaccination coverage rates (0.87%-18%) and an overall decrease from 26% to 19%, highlighting the impact of missing data on estimates.
Conclusions:
- MICE and miceforest provide reliable methods for imputing missing demographic data, mitigating bias more effectively than Iterative-Imputer. Miceforest offers enhanced computational efficiency, especially for large datasets, through cloud-based processing.
- The choice of imputation method significantly impacts research findings, underscoring the need for careful selection.
- The substantial decrease in estimated vaccination coverage emphasizes how missing data can obscure true disparities. Regular application of imputation methods is recommended to improve health equity evaluations and guide targeted public health interventions.
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Statistical Methods for Analyzing Epidemiological Data
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Bias in Epidemiological Studies
Immunodeficiency Diseases
There are three main causes of immunodeficiency...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...

