Related Experiment Video
Updated: Apr 26, 2026

Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
Methods for mediation analysis with missing data
1University of Notre Dame, Notre Dame, IN, USA, zzhang4@nd.edu.
Abstract:
Despite wide applications of both mediation models and missing data techniques, formal discussion of mediation analysis with missing data is still rare. We introduce and compare four approaches to dealing with missing data in mediation analysis including listwise deletion, pairwise deletion, multiple imputation (MI), and a two-stage maximum likelihood (TS-ML) method. An R package bmem is developed to implement the four methods for mediation analysis with missing data in the structural equation modeling framework, and two real examples are used to illustrate the application of the four methods. The four methods are evaluated and compared under MCAR, MAR, and MNAR missing data mechanisms through simulation studies. Both MI and TS-ML perform well for MCAR and MAR data regardless of the inclusion of auxiliary variables and for AV-MNAR data with auxiliary variables. Although listwise deletion and pairwise deletion have low power and large parameter estimation bias in many studied conditions, they may provide useful information for exploring missing mechanisms.
Insights
This study compares four methods for mediation analysis with missing data: listwise deletion, pairwise deletion, multiple imputation (MI), and two-stage maximum likelihood (TS-ML). MI and TS-ML are recommended for most missing data scenarios.
Area of Science:
- Statistics
- Psychometrics
- Quantitative Psychology
Background:
- Mediation analysis is widely used but formally addressing missing data is uncommon.
- Existing methods for missing data in mediation analysis lack comprehensive comparison.
Purpose of the Study:
- To introduce and compare four methods for handling missing data in mediation analysis.
- To evaluate the performance of these methods under different missing data mechanisms.
- To provide practical tools for mediation analysis with missing data.
Main Methods:
- Comparison of listwise deletion, pairwise deletion, multiple imputation (MI), and two-stage maximum likelihood (TS-ML).
- Development of the R package 'bmem' for implementing these methods within structural equation modeling.
- Simulation studies under Missing Completely At Random (MCAR), Missing At Random (MAR), and Missing Not At Random (MNAR) conditions.
Main Results:
- Multiple imputation (MI) and two-stage maximum likelihood (TS-ML) demonstrated strong performance for MCAR and MAR data.
- MI and TS-ML were effective for Auxiliary Variable Missing Not At Random (AV-MNAR) data when auxiliary variables were included.
- Listwise and pairwise deletion showed low statistical power and significant parameter estimation bias in many scenarios.
Conclusions:
- Multiple imputation (MI) and two-stage maximum likelihood (TS-ML) are robust methods for mediation analysis with missing data.
- The R package 'bmem' facilitates the application of these advanced techniques.
- While less powerful, deletion methods can offer insights into missing data mechanisms.
More Related Videos
Related Concept Videos
Statistical Methods to Analyze Parametric Data: ANOVA
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
One-Way ANOVA: Unequal Sample Sizes
Mechanistic Models: Compartment Models in Individual and Population Analysis
Censoring Survival Data
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...

