Related Experiment Video
Updated: Jun 12, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
The performance of prognostic models depended on the choice of missing value imputation algorithm: a simulation study
Manja Deforth1, Georg Heinze2, Ulrike Held1
1Department of Biostatistics at the Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Zurich, Switzerland.
Multiple imputation methods like mice and aregImpute are reliable choices for handling missing predictor values in clinical prediction models, offering good performance even with complex data. These methods improve model calibration and discrimination when data is incomplete.
Area of Science:
- Biostatistics
- Machine Learning in Healthcare
- Clinical Prediction Modeling
Background:
- Missing predictor values frequently hinder the development of robust clinical prediction models.
- Existing imputation methods vary in complexity, from simple single imputation to advanced multiple imputation techniques using chained equations.
- Machine learning algorithms and flexible modeling are increasingly integrated into imputation strategies.
Purpose of the Study:
- To evaluate the comparative performance of different missing value imputation methods in the context of clinical prediction model development.
- To determine if specific imputation algorithms consistently outperform others across various performance metrics.
Main Methods:
- Simulated development and validation cohorts mimicking real data distributions.
- Applied three R-based imputation algorithms: mice, aregImpute, and missForest, under 36 missingness scenarios.
- Assessed model performance using Brier score, c-statistic, calibration, and prediction error, comparing against full data analysis.
Main Results:
- No imputation method fully replicated the performance of complete data; complete case analysis performed worst.
- aregImpute and mice (with 100 imputations) demonstrated the highest predictive accuracy across scenarios.
- aregImpute notably achieved calibration slopes close to one, outperforming even full data analysis in this aspect.
Conclusions:
- Model calibration is more sensitive to imputation method choice than discrimination.
- Multiple imputation methods, particularly mice and aregImpute, are robust and recommended for handling missing data in prediction models.
- These methods effectively manage linear and nonlinear predictor-outcome associations, providing reliable results.
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Kaplan-Meier Approach
Censoring Survival Data
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Comparing the Survival Analysis of Two or More Groups

