Related Experiment Video
Updated: Oct 12, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
[Simulation study on missing data imputation methods for longitudinal data in cohort studies]
1Department of Epidemiology and Biostatistics, School of Public Health of Xi'an Jiaotong University Health Science Center, Xi'an 710061, China.
For longitudinal studies, mean imputation, k-nearest neighbor (KNN), regression imputation, and random forest are effective for handling missing data. Other methods like K-means clustering and expectation maximization (EM) are not recommended due to instability.
Area of Science:
- Biostatistics
- Epidemiology
- Data Science
Background:
- Missing data is a common challenge in longitudinal cohort studies.
- Effective imputation methods are crucial for reliable data analysis and subsequent statistical modeling.
Purpose of the Study:
- To compare the performance of eight common missing data imputation techniques.
- To evaluate their impact on longitudinal data and subsequent multivariate analyses.
- To provide guidance for selecting appropriate imputation methods in cohort studies.
Main Methods:
- A simulation study was conducted using R language software.
- Longitudinal data with missing values were generated using the Monte Carlo method.
- Imputation methods were evaluated based on average absolute deviation, average relative deviation, and Type I error in regression analysis.
Main Results:
- Mean imputation, k-nearest neighbor (KNN), regression imputation, and random forest demonstrated stable and comparable imputation effects.
- Hot deck imputation performed less effectively than the aforementioned methods.
- K-means clustering and expectation maximization (EM) algorithm showed the poorest and most unstable results.
- Mean imputation, EM algorithm, random forest, KNN, and regression imputation effectively controlled Type I error, while multiple imputations, hot deck, and K-means clustering did not.
Conclusions:
- Mean imputation, KNN, regression imputation, and random forest are recommended for missing data in longitudinal studies under the missing at random mechanism.
- Multiple imputations and hot deck can be suitable when the missing data ratio is low.
- K-means clustering and EM algorithm are not advised due to their instability and poor performance.
Related Concept Videos
Longitudinal Studies
Longitudinal Research
Censoring Survival Data
Assumptions of Survival Analysis
Comparing the Survival Analysis of Two or More Groups
Kaplan-Meier Approach

