Related Experiment Video
Updated: Jul 15, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A blocking strategy to improve gene selection for classification of gene expression data.
1Départment d'Informative,Université Libre de Bruxelles, Bruxelles, Belgium. gbonte@ulb.ac.be
This study introduces a novel blocking strategy for gene feature selection in microarray data. This method enhances classification accuracy by robustly evaluating gene subsets across multiple learning algorithms, improving reliability in high-dimensional datasets.
Area of Science:
- Bioinformatics
- Machine Learning
- Computational Biology
Background:
- High-dimensional microarray gene expression data presents challenges for machine learning due to the large number of features relative to samples.
- Traditional feature selection methods are computationally intensive and prone to errors in such datasets.
- Feature selection is crucial for effective classification and requires robust evaluation of gene subsets.
Purpose of the Study:
- To propose an original blocking strategy to improve feature selection in high-dimensional microarray gene expression data.
- To enhance the reliability and accuracy of gene subset selection by aggregating validation outcomes from multiple learning algorithms.
- To address the computational hardness and error-proneness of feature selection in datasets with many features and few samples.
Main Methods:
- Interpreting feature selection as a stochastic optimization task to find gene subsets with optimal classification generalization.
- Developing a novel blocking strategy that pairs validation outcomes from multiple learning algorithms to assess gene subsets.
- Comparing the proposed blocking strategy against conventional wrapper methods using 16 cancer expression datasets and six different classifiers.
Main Results:
- The proposed blocking strategy significantly improved the performance of conventional forward selection methods.
- Improvements in feature selection accuracy were observed independently of the classification algorithm used post-selection.
- Biological validation using PubMed abstracts and Gene Ontology analysis supported the enhanced accuracy and quality of selected gene subsets.
Conclusions:
- The developed blocking strategy offers a more robust and accurate approach to feature selection in high-dimensional gene expression data.
- Aggregating validation from multiple algorithms via blocking mitigates issues arising from limited sample sizes.
- This method holds promise for improving the discovery of biologically relevant genes from expression datasets.
More Related Videos
03:08Using Human Differentially Expressed Gene Lists to Perform Downstream Pathway Enrichment Analysis and Target Prioritization
Published on: October 3, 2025
09:35A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017