Related Experiment Video
Updated: May 17, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Combining multiple hypothesis testing and affinity propagation clustering leads to accurate, robust and sample size
Argiris Sakellariou1, Despina Sanoudou, George Spyrou
1Biomedical Informatics Unit, Biomedical Research Foundation of the Academy of Athens, Athens, Greece.
BMC Bioinformatics
|October 19, 2012
Summary
A new hybrid feature selection method (mAP-KL) identifies biologically relevant gene clusters for disease classification. This approach improves diagnostic and prognostic accuracy across diverse microarray datasets.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Feature selection in gene expression data needs to be platform, disease, and dataset size independent.
- Hypothesis: Statistically significant genes contain clusters with shared biological functions relevant to disease.
- Proposes using gene cluster exemplars instead of a fixed number of top-ranked genes.
Purpose of the Study:
- To develop and evaluate a hybrid feature selection method (mAP-KL) for identifying informative gene subsets from microarray data.
- To improve the accuracy and biological relevance of gene signatures for disease classification.
- To create a data-driven, classifier-independent method applicable to various diseases and dataset sizes.
Main Methods:
- Combines multiple hypothesis testing with affinity propagation (AP) clustering.
- Utilizes the Krzanowski & Lai cluster quality index.
- Applied to real and simulated microarray data for performance comparison.
Main Results:
- mAP-KL demonstrated competitive classification results across various diseases and sample sizes.
- Achieved a high Area Under the Curve (AUC) score of 0.91 in neuromuscular diseases.
- Generated concise, biologically relevant gene expression signatures for diagnostic/prognostic use.
Conclusions:
- mAP-KL is a data-driven, classifier-independent hybrid feature selection method.
- Applicable to any disease classification problem using microarray data, irrespective of sample size.
- Effectively classifies samples from both small and large cohorts with high accuracy.
Related Concept Videos
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Accuracy and Errors in Hypothesis Testing
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Statistical Hypothesis Testing
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
