Related Experiment Video
Updated: Sep 24, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Partial Multiple Imputation With Variational Autoencoders: Tackling Not at Randomness in Healthcare Data
Abstract:
Missing data can pose severe consequences in critical contexts, such as clinical research based on routinely collected healthcare data. This issue is usually handled with imputation strategies, but these tend to produce poor and biased results under the Missing Not At Random (MNAR) mechanism. A recent trend that has been showing promising results for MNAR is the use of generative models, particularly Variational Autoencoders. However, they have a limitation: the imputed values are the result of a single sample, which can be biased. To tackle it, an extension to the Variational Autoencoder that uses a partial multiple imputation procedure is introduced in this work. The proposed method was compared to 8 state-of-the-art imputation strategies, in an experimental setup with 34 datasets from the medical context, injected with the MNAR mechanism (10% to 80% rates). The results were evaluated through the Mean Absolute Error, with the new method being the overall best in 71% of the datasets, significantly outperforming the remaining ones, particularly for high missing rates. Finally, a case study of a classification task with heart failure data was also conducted, where this method induced improvements in 50% of the classifiers.
Insights
This study introduces an improved Variational Autoencoder method for handling missing healthcare data, outperforming existing strategies in 71% of medical datasets, especially with high missing rates.
Area of Science:
- Data Science
- Machine Learning
- Biostatistics
Background:
- Missing data in healthcare research, especially from routinely collected data, can lead to biased results.
- Traditional imputation methods often fail under Missing Not At Random (MNAR) mechanisms.
- Generative models like Variational Autoencoders show promise but can produce biased imputations from single samples.
Purpose of the Study:
- To develop an advanced imputation method addressing limitations of standard Variational Autoencoders for MNAR data.
- To enhance the accuracy and reduce bias in imputing missing values in critical healthcare datasets.
- To improve the performance of machine learning models using data with missing values.
Main Methods:
- An extension of Variational Autoencoders incorporating a partial multiple imputation procedure was developed.
- The proposed method was evaluated against 8 state-of-the-art imputation strategies.
- Experiments were conducted on 34 medical datasets with simulated MNAR data at missing rates from 10% to 80%.
Main Results:
- The novel method achieved superior performance, outperforming existing strategies in 71% of the evaluated datasets.
- Significant improvements were observed, particularly at higher missing data rates (up to 80%).
- In a heart failure classification case study, the method improved the performance of 50% of the tested classifiers.
Conclusions:
- The proposed partial multiple imputation extension to Variational Autoencoders offers a robust solution for MNAR data in healthcare.
- This method significantly enhances imputation accuracy and reduces bias compared to current state-of-the-art techniques.
- The approach demonstrates potential for improving downstream machine learning tasks in clinical research.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Statistical Methods for Analyzing Epidemiological Data
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...

