Related Experiment Video
Updated: Jun 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Optimal classifier selection and negative bias in error rate estimation: an empirical study on high-dimensional
Anne-Laure Boulesteix1, Carolin Strobl
1Department of Statistics, University of Munich, Ludwigstr 33, D-80539 Munich, Germany. boulesteix@ibe.med.uni-muenchen.de
Researchers often select the best results from many tests, leading to biased error estimates in high-dimensional data analysis. This practice inflates accuracy and is not acceptable for reliable biometric predictions.
Area of Science:
- Biometrics
- Machine Learning
- Bioinformatics
Background:
- Biometric research frequently employs a "trial-and-error" approach, testing numerous methods to maximize data utility.
- Publication and client pressures can lead to the selective reporting of only the most favorable outcomes.
- This selective reporting induces significant optimistic bias in prediction error estimation, particularly with high-dimensional data like microarrays.
Purpose of the Study:
- To quantitatively assess the optimistic bias in prediction error estimation caused by data-driven classifier selection.
- To investigate the bias arising from parameter tuning, variable selection, and classifier choice in high-dimensional data analysis.
- To evaluate the impact of these biases on the reliability of classification accuracy in biometrics.
Main Methods:
- Examined 124 classifier variants, including those with variable selection and parameter tuning, within a cross-validation framework.
- Applied classifiers to real and modified microarray datasets, including those with randomly permuted class labels to simulate non-informative predictors.
- Assessed the minimal misclassification rate across classifier variants to quantify selection bias.
Main Results:
- The minimal misclassification rate was evaluated to quantify bias from data-driven optimal classifier selection.
- Bias from parameter tuning (including gene selection) and classifier method choice were analyzed separately and jointly.
- Median minimal error rates were 31% (colon cancer) and 41% (prostate cancer) on permuted data, indicating substantial bias.
Conclusions:
- The strategy of reporting only the optimal result introduces substantial bias in error rate estimation.
- This selective reporting is unacceptable for accurate biometric predictions.
- Alternative methods for robustly reporting classification accuracy are recommended to mitigate bias.
Related Concept Videos
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies
Unrealistic Optimism Bias
Errors In Hypothesis Tests
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...