Related Experiment Video
Updated: Jun 2, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Multiple-rule bias in the comparison of classification rules
Mohammadmahdi R Yousefi1, Jianping Hua, Edward R Dougherty
1Department of Electrical and Computer Engineering, Texas A&M University, College Station, TX 77843, USA.
This study addresses overoptimism in bioinformatics classification by analyzing the "multiple-rule bias." It quantifies how selecting the best-performing rule on a dataset can inflate performance estimates, impacting reported results.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Learning
Background:
- Bioinformatics community faces overoptimism in reported classification results.
- Two key issues contributing to overoptimism include reporting on favorable datasets and comparing multiple rules on a single dataset.
Purpose of the Study:
- To provide a probabilistic analysis of the 'multiple-rule bias' in classification.
- To quantify the bias in estimating the true error of the selected classification rule.
- To characterize the bias in estimating the comparative advantage of the chosen rule.
Main Methods:
- Probabilistic analysis of classification rule selection bias.
- Application of analysis to synthetic and real datasets.
- Utilizing various classification rules and error estimation methods.
Main Results:
- Quantification of the bias when selecting a classification rule with minimum estimated error.
- Characterization of bias in estimating the true comparative advantage of the selected rule.
- Demonstration of bias using both synthetic and real-world data.
Conclusions:
- The 'multiple-rule bias' is a significant factor contributing to overoptimism in classification performance reporting.
- Accurate estimation of classification rule performance requires accounting for the bias introduced by rule selection.
- The study provides a framework for more reliable evaluation of classification methods in bioinformatics.
Related Concept Videos
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
The Representativeness Heuristic
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Bias in Epidemiological Studies
Motivational Bias
