Related Experiment Video
Updated: Sep 3, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Multiple imputation and test-wise deletion for causal discovery with incomplete cohort data.
Janine Witte1,2, Ronja Foraita1, Vanessa Didelez1,2
1Leibniz Institute for Prevention Research and Epidemiology - BIPS, Bremen, Germany.
Causal discovery algorithms can now handle missing data. Test-wise deletion and multiple imputation outperform older methods, with multiple imputation being particularly effective for small, homogenous datasets.
Area of Science:
- Causal inference and machine learning
- Statistical modeling and data analysis
Background:
- Causal discovery algorithms estimate causal graphs from observational data, complementing individual treatment-outcome analyses.
- Constraint-based causal discovery algorithms traditionally struggle with missing data, limiting their application.
Purpose of the Study:
- To investigate and compare the performance of test-wise deletion and multiple imputation for handling missing values in causal discovery.
- To establish conditions for causal structure recoverability using test-wise deletion.
Main Methods:
- Simulated data from benchmark causal graphs were used to compare various imputation methods.
- Methods evaluated include test-wise deletion, multiple imputation, list-wise deletion, single imputation, random forest imputation, and a hybrid approach.
Main Results:
- Both test-wise deletion and multiple imputation significantly outperform list-wise deletion and single imputation.
- Multiple imputation shows particular utility with small datasets containing only Gaussian or discrete variables.
- Performance is mixed when datasets contain a combination of Gaussian and discrete variables.
Conclusions:
- Test-wise deletion and multiple imputation offer viable solutions for incorporating missing data into causal discovery.
- The choice of method depends on the characteristics of the data, particularly variable types and dataset size.
Related Concept Videos
Censoring Survival Data
Causality in Epidemiology
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Comparing the Survival Analysis of Two or More Groups
Assumptions of Survival Analysis
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...

