Related Experiment Video
Updated: Oct 12, 2025

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Missing data is poorly handled and reported in prediction model studies using machine learning: a literature review
Swj Nijman1, A M Leeuwenberg1, I Beekers2
1Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht University, Heidelberglaan 100, Utrecht, 3584 CX , The Netherlands.
Many machine learning prediction model studies fail to adequately report missing data handling. Deletion methods are common, despite potential bias, highlighting a need for better reporting and alternative strategies in medical research.
Area of Science:
- Medical research
- Machine learning
- Prediction models
Background:
- Missing data is a pervasive challenge in developing and evaluating prediction models.
- Machine learning (ML) methods are often perceived as capable of handling missing data, but their application in medical research lacks clarity.
- There is a need to assess how ML-based prediction model studies report their strategies for managing missing data.
Purpose of the Study:
- To investigate the extent and methods of missing data reporting in ML-based clinical prediction model studies.
- To evaluate the quality of reporting regarding the amount, nature, and handling of missing data in these studies.
Main Methods:
- Systematic literature search of primary studies published between 2018-2019.
- Inclusion of studies developing or validating clinical prediction models using supervised ML methodologies.
- Extraction of data on missing data presence, characteristics, and handling methods.
Main Results:
- 152 ML-based clinical prediction model studies were identified.
- 56% (56/152) of studies did not report on missing data.
- Of those reporting, many failed to specify the amount of missingness (46/96), reasons for missingness (7/96), or data mechanisms (8/96).
- Deletion methods, particularly complete-case analysis (43/96), were the most common approach (65/96).
- Multiple imputation (8/96) and built-in ML mechanisms (7/96) were rarely used.
Conclusions:
- A majority of ML prediction model studies provide insufficient information on missing data.
- Commonly used deletion strategies can introduce bias and reduce analytical power.
- Researchers should improve reporting transparency and consider advanced methodologies for handling missing data.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Regression Toward the Mean
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Steps in Outbreak Investigation
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...

