Related Experiment Video
Updated: Apr 30, 2026

MicroRNA Based Liquid Biopsy: The Experience of the Plasma miRNA Signature Classifier MSC for Lung Cancer Screening
Published on: October 26, 2017
ML versus MI for Missing Data with Violation of Distribution Conditions
Ke-Hai Yuan1, Fan Yang-Wallentin2, Peter M Bentler3
1University of Notre Dame.
Maximum likelihood (ML) is generally preferable to multiple imputation (MI) for missing data analysis. ML offers more efficient parameter estimates and reliable standard errors, especially with non-normal data.
Area of Science:
- Statistics
- Data Analysis
Background:
- Missing data analysis is crucial in statistical modeling.
- Maximum Likelihood (ML) and Multiple Imputation (MI) are primary methods.
- Comparing their performance regarding bias and efficiency is essential.
Purpose of the Study:
- To compare ML and MI for missing data analysis.
- To evaluate bias and efficiency of parameter estimates.
- To assess standard error estimation accuracy.
Main Methods:
- Comparative analysis of ML and MI procedures.
- Evaluation of parameter estimate bias and efficiency.
- Comparison of formula-based versus empirical standard errors.
Main Results:
- MI parameter estimates are less efficient than ML.
- MI variance-covariance estimates show more bias, especially with heavy-tailed distributions.
- ML estimates can be biased at smaller sample sizes.
- Sandwich-type and observed information matrix SEs are comparable to empirical SEs for normal distributions.
- Sandwich-type SEs are more reliable for ML with heavy-tailed distributions.
- Neither formula-based SE method is consistent for MI.
Conclusions:
- ML is generally preferred over MI in practice for missing data.
- MI parameter estimates may still be consistent despite limitations.
- Accurate standard error estimation is critical and challenging with MI.
Related Concept Videos
One-Way ANOVA: Unequal Sample Sizes
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Modified Boxplots
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
Mechanistic Models: Compartment Models in Individual and Population Analysis
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...

