Related Experiment Video
Updated: Jul 4, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Classification of genomic data: some aspects of feature selection
Tomasz Czekaj1, Wen Wu, Beata Walczak
1Department of Chemometrics, Institute of Chemistry, The University of Silesia, 9 Szkolna Street, 40-006 Katowice, Poland.
Selecting relevant genes from genomic data improves classification and interpretability. However, different feature selection methods yield varying gene subsets, requiring careful interpretation due to gene correlations.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Feature selection is crucial for genomic data analysis, enhancing both diagnostic classification and data interpretability.
- Multivariate approaches effectively reduce data dimensionality but can lead to variable subset dependency on classifier objectives.
Purpose of the Study:
- To investigate the impact of different objective functions on feature selection outcomes in genomic datasets.
- To highlight the importance of cautious interpretation of selected features due to potential gene correlations.
Main Methods:
- Applied multivariate feature selection techniques to genomic datasets.
- Evaluated the dependency of selected feature subsets on various classifier objective functions.
- Analyzed gene correlations within selected subsets.
Main Results:
- Demonstrated that the composition of selected gene subsets is contingent upon the specific objective function employed by the classifier.
- Identified that correlated genes can lead to the selection of different subsets for classification tasks.
- Confirmed the necessity of careful interpretation of selected gene sets.
Conclusions:
- The choice of objective function significantly influences feature selection in genomic data.
- Understanding gene correlations is vital for accurate interpretation of classification-related feature subsets.
- Feature selection strategies must consider both predictive power and interpretability in genomic studies.
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Evolutionary Relationships through Genome Comparisons
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which result in visible changes...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantifying and Rejecting Outliers: The Grubbs Test
