Related Experiment Video
Updated: Jun 15, 2026

In Vivo Modeling of the Morbid Human Genome using Danio rerio
Published on: August 24, 2013
Impact of missing value imputation on classification for DNA microarray gene expression data--a model-based study
Youting Sun1, Ulisses Braga-Neto, Edward R Dougherty
1Department of Electrical and Computer Engineering, Texas A&M University, College Station, TX 77843, USA.
Abstract:
Many missing-value (MV) imputation methods have been developed for microarray data, but only a few studies have investigated the relationship between MV imputation and classification accuracy. Furthermore, these studies are problematic in fundamental steps such as MV generation and classifier error estimation. In this work, we carry out a model-based study that addresses some of the issues in previous studies. Six popular imputation algorithms, two feature selection methods, and three classification rules are considered. The results suggest that it is beneficial to apply MV imputation when the noise level is high, variance is small, or gene-cluster correlation is strong, under small to moderate MV rates. In these cases, if data quality metrics are available, then it may be helpful to consider the data point with poor quality as missing and apply one of the most robust imputation algorithms to estimate the true signal based on the available high-quality data points. However, at large MV rates, we conclude that imputation methods are not recommended. Regarding the MV rate, our results indicate the presence of a peaking phenomenon: performance of imputation methods actually improves initially as the MV rate increases, but after an optimum point, performance quickly deteriorates with increasing MV rates.
Insights
Missing value imputation in microarray data can improve classification accuracy under certain conditions, especially with high noise or strong correlations. However, imputation is not recommended for large amounts of missing data.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Numerous missing-value (MV) imputation methods exist for microarray data.
- Few studies have rigorously examined the impact of MV imputation on classification accuracy.
- Previous research often suffers from flawed MV generation and error estimation.
Purpose of the Study:
- To conduct a model-based investigation into the relationship between MV imputation and classification accuracy in microarray data.
- To address fundamental issues identified in prior studies concerning MV handling and evaluation.
Main Methods:
- Evaluation of six popular MV imputation algorithms.
- Incorporation of two feature selection methods.
- Assessment using three distinct classification rules.
- Model-based simulation to control for MV generation and error estimation.
Main Results:
- MV imputation is beneficial when noise levels are high, variance is low, or gene-cluster correlations are strong, particularly at small to moderate MV rates.
- Data quality metrics can guide imputation by identifying poor-quality points for imputation.
- Performance of imputation methods exhibits a peaking phenomenon, improving up to an optimal MV rate before deteriorating rapidly.
Conclusions:
- MV imputation is recommended for specific data characteristics and moderate missingness rates.
- Imputation methods are not advised for high missing value rates.
- The optimal MV rate for imputation performance should be considered to avoid accuracy degradation.

