Related Experiment Video
Updated: May 2, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Comparison of methods to handle missing values in a binary index test in a diagnostic accuracy study - a simulation
Dennis Juljugin1, Katharina Stahlmann2, Antonia Zapf1
1Institute of Medical Biometry and Epidemiology, University Medical Center Hamburg-Eppendorf, Hamburg, Germany.
Background:
As there are no recommendations on handling missing values in a dichotomous index test of a diagnostic study, researchers often ignore missing values in the analysis or use simple methods. Thus, this simulation study compares selected methods regarding their performance of estimating sensitivity and specificity of a dichotomous index test with missing values.
Methods:
Data of a single-test diagnostic study were simulated including a dichotomous reference standard, a dichotomous index test and three dichotomous covariates. Following different proportions of missing values and missingness mechanisms, missing values were modeled in the index test. Additionally, the sample size, true sensitivity and specificity, and the prevalence of the target condition were varied in the data generation resulting in 729 scenarios. Seven methods were compared: complete case analysis, worst case scenario (WC), random hot deck, multiple imputation by chained equations (MICE), and three different product multinomial framework approaches.
Results:
Apart from WC, most methods are unbiased under missing completely at random (MCAR). Under missing at random (MAR), however, MICE clearly outperforms the other methods and is nearly unbiased while the other methods are considerably more biased. Additionally, MICE shows the best coverage probability for MCAR and MAR. If missing values are missing not at random (MNAR), all methods are substantially biased and show coverage probability that is too low.
Conclusions:
While most methods perform well when the proportion of missing values is small, especially under MCAR, MICE should be used when the proportion of missing values increases and the missing values are MAR. None of the tested methods seems to be suitable for MNAR.
Related Concept Videos
Kaplan-Meier Approach
Comparing the Survival Analysis of Two or More Groups
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Receiver Operating Characteristic Plot
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Censoring Survival Data

