Related Experiment Video
Updated: Oct 26, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Evaluating the impact of multivariate imputation by MICE in feature selection
Maritza Mera-Gaona1, Ursula Neumann2, Rubiel Vargas-Canas1
1University of Cauca, Colombia, Popayán, Cauca, Colombia.
Multivariate imputation significantly improves machine learning feature selection on datasets with missing values. This method, unlike basic imputation techniques, reduces bias and enhances analytical accuracy.
Area of Science:
- Machine Learning
- Data Preprocessing
- Statistical Analysis
Background:
- Handling missing values is critical in machine learning.
- Traditional methods like discarding data or using mean/median imputation can introduce bias.
- Most machine learning algorithms require complete datasets for analysis.
Purpose of the Study:
- To demonstrate the positive impact of multivariate imputation on feature selection for datasets with missing values.
- To compare multivariate imputation against basic imputation techniques and complete case analysis.
Main Methods:
- Feature selection was performed on complete datasets, incomplete datasets (5-50% missingness), and datasets imputed using basic techniques and multivariate imputation (MICE).
- Standard feature selection algorithms were employed for comparison.
Main Results:
- Datasets imputed using multivariate imputation (MICE) yielded superior results in the feature selection process.
- Multivariate imputation outperformed basic imputation methods and non-imputed incomplete datasets.
Conclusions:
- Applying multivariate imputation, specifically MICE, effectively reduces bias in the feature selection process.
- Multivariate imputation is a recommended strategy for handling missing data in machine learning workflows.
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Comparing the Survival Analysis of Two or More Groups
One-Way ANOVA: Unequal Sample Sizes
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

