Related Experiment Video
Updated: Oct 23, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Multiple imputation of semi-continuous exposure variables that are categorized for analysis
Cattram D Nguyen1,2, Margarita Moreno-Betancur1,2, Laura Rodwell1,2,3
1Clinical Epidemiology and Biostatistics Unit, Murdoch Children's Research Institute, Parkville, Victoria, Australia.
Choosing the right method for imputing semi-continuous data, like alcohol consumption, is crucial. Ordinal logistic regression and zero-inflated binomial imputation performed best for handling missing values and subsequent categorization.
Area of Science:
- Biostatistics
- Statistical Modeling
- Epidemiology
Background:
- Semi-continuous variables, common in health research (e.g., alcohol consumption), possess a unique distribution with a point mass at zero and continuous positive values.
- Handling missing data in these variables using multiple imputation presents challenges, particularly when analysis requires categorized exposure variables.
Purpose of the Study:
- To compare the performance of various multiple imputation methods for semi-continuous exposure variables that require categorization for analysis.
- To assess the congeniality between imputation models and analysis models for semi-continuous data.
Main Methods:
- A simulation study evaluated nine imputation approaches for semi-continuous exposures.
- Methods included direct imputation of categories (ordinal logistic regression, multivariate normal imputation [MVNI] of indicator variables) and imputation of the continuous variable followed by categorization (predictive mean matching, zero-inflated binomial imputation, two-part methods).
- Performance was assessed based on different estimands (proportions, regression coefficients).
Main Results:
- Ordinal logistic regression and zero-inflated binomial imputation demonstrated good performance across most simulation scenarios.
- Multivariate normal imputation (MVNI) methods that required rounding after imputation performed poorly.
- Predictive mean matching and two-part methods showed mixed results, dependent on the specific estimand.
Conclusions:
- The choice of imputation method for semi-continuous variables significantly impacts analysis results, especially when categorization is required.
- Ordinal logistic regression and zero-inflated binomial imputation are recommended for their robust performance in these scenarios.
- Researchers must carefully consider the parameter of interest when selecting an imputation procedure for semi-continuous data.
More Related Videos
09:50Impact Assessment of Repeated Exposure of Organotypic 3D Bronchial and Nasal Tissue Culture Models to Whole Cigarette Smoke
Published on: February 12, 2015
10:46A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Biostatistics: Overview
Discrete variables are...
Censoring Survival Data
Mechanistic Models: Compartment Models in Individual and Population Analysis
Statistical Methods for Analyzing Epidemiological Data
Comparing the Survival Analysis of Two or More Groups