Related Experiment Video
Updated: Feb 3, 2026

Pattern-based Search of Epigenomic Data Using GeNemo
Published on: October 8, 2017
Classifying Incomplete Gene-Expression Data: Ensemble Learning with Non-Pre-Imputation Feature Filtering and
Yuanting Yan1,2, Tao Dai3, Meili Yang4
1School of Computer Science and Technology, Anhui University, Hefei 230601, China. ytyan2016@163.com.
This study introduces a novel feature selection method for gene expression data that directly handles missing values, bypassing imputation. This approach improves classification accuracy by avoiding imputation-induced biases in gene selection.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Gene-expression datasets frequently contain missing values (MVs).
- While MV imputation methods are abundant, their impact on downstream classification performance is minimal.
- Prioritizing informative gene selection over MV imputation is crucial for classification tasks.
Purpose of the Study:
- To develop a feature selection (FS) method that directly processes incomplete gene-expression data without imputation.
- To evaluate the impact of MV imputation on downstream FS performance.
- To identify informative genes for classification from gene-expression data with missing values.
Main Methods:
- A modified chi-square test-based FS approach is proposed for gene-expression data.
- Recursive element aggregation is introduced to address small sample sizes in gene-expression data.
- The method directly handles incomplete data, avoiding imputation and its assumptions, followed by best-first search for optimal feature subsets.
Main Results:
- The proposed method was compared with existing FS algorithms on twelve incomplete cancer gene-expression datasets.
- MV imputation can negatively impact subsequent FS and classification performance.
- Directly applying FS to incomplete data avoids imputation-related disturbances, leading to potentially better feature subsets.
Conclusions:
- The developed FS method effectively handles missing values in gene-expression data without imputation.
- This approach offers a more robust alternative to traditional FS methods that require prior imputation.
- The method demonstrated improved gene discovery by identifying additional informative genes on the SRBCT dataset.
More Related Videos
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
10:34Using an Automated Cell Counter to Simplify Gene Expression Studies: siRNA Knockdown of IL-4 Dependent Gene Expression in Namalwa Cells
Published on: April 14, 2010
Related Concept Videos
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
What is Gene Expression?
Gene expression is the process in which DNA directs the synthesis of functional products, that is, proteins. Cells can regulate gene expression at various stages. It allows organisms to generate different cell types and enables cells to adapt to internal and external factors.
Genetic Information Flows from DNA to RNA to Protein
A gene is a stretch of DNA that serves as the blueprint for functional RNAs and proteins. Since DNA is made up of nucleotides and proteins consist of amino...
What is Gene Expression?
Cell Specific Gene Expression
Chromatin Position Affects Gene Expression
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...