Related Experiment Video
Updated: Jan 12, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Evaluation of missing data analytical techniques in longitudinal research: Traditional and machine learning
1Department of Psychology, University of Virginia.
Full information maximum likelihood (FIML) is most effective for missing not at random (MNAR) data in growth curve models. Two-stage robust estimation (TSRE) handles missing at random (MAR) data well, while machine learning methods show limited benefits.
Area of Science:
- Statistical modeling
- Longitudinal data analysis
- Machine learning applications
Background:
- Handling nonnormal and missing not at random (MNAR) data is complex for traditional methods like full information maximum likelihood (FIML) due to normal distribution assumptions.
- Two-stage robust estimation (TSRE) can handle nonnormal data, but its performance in longitudinal studies with MNAR conditions is less explored.
- Machine learning (ML) offers an alternative, not requiring distributional assumptions and showing promise for MNAR data, though its use in longitudinal studies for both missing at random (MAR) and MNAR is underexplored.
Purpose of the Study:
- To compare the effectiveness of six analytical techniques for missing data within growth curve modeling.
- To assess the impact of sample size, missing data rate, mechanism, and distribution on model estimation accuracy and efficiency.
- To evaluate traditional (FIML, TSRE) and machine learning (K-nearest neighbors, missForest, micecart, miceForest) imputation methods.
Main Methods:
- Monte Carlo simulations were employed to evaluate analytical techniques.
- Growth curve modeling framework was used to assess missing data handling.
- Six techniques were compared: FIML, TSRE, K-nearest neighbors, missForest, micecart, and miceForest.
Main Results:
- Full information maximum likelihood (FIML) demonstrated the highest effectiveness for missing not at random (MNAR) data.
- Two-stage robust estimation (TSRE) performed best for missing at random (MAR) data.
- MissForest showed advantages only under specific conditions: highly skewed distributions, large sample sizes (n ≥ 1,000), and low missing data rates.
Conclusions:
- FIML is recommended for longitudinal growth curve models with MNAR data.
- TSRE is a suitable choice for MAR data scenarios.
- Machine learning imputation methods, like missForest, have niche applications but require careful consideration of data characteristics and sample size.
Related Concept Videos
Longitudinal Studies
Longitudinal Research
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Steps in Outbreak Investigation
Kaplan-Meier Approach

