Related Experiment Video
Updated: Aug 14, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Feature selection and classifier performance in computer-aided diagnosis: the effect of finite sample size
B Sahiner1, H P Chan, N Petrick
1Department of Radiology, University of Michigan, Ann Arbor 48109-0904, USA. berki@umich.edu
Medical Physics
|August 18, 2000
Summary
Finite sample size significantly impacts classification accuracy in computer-aided diagnosis (CAD) using stepwise feature selection. Hold-out estimates are biased when feature selection uses only design samples, but can be biased either way when using all samples.
Area of Science:
- Medical Imaging
- Machine Learning
- Statistical Analysis
Background:
- Computer-aided diagnosis (CAD) often involves feature extraction and selection for classification.
- Stepwise feature selection is common for linear classifiers, but its performance with finite sample sizes is not fully understood.
- Classifier design stages, including feature selection and coefficient estimation, can interact and affect accuracy.
Purpose of the Study:
- To investigate the impact of finite sample size on classification accuracy in CAD.
- To analyze the bias in performance estimates (hold-out and resubstitution) during stepwise feature selection.
- To compare feature selection performed on design samples versus the entire sample set.
Main Methods:
- Simulated classification tasks with multidimensional Gaussian distributions and a large number of features.
- Employed stepwise feature selection within linear discriminant analysis.
- Evaluated classifier performance using the area Az under the receiver operating characteristic curve.
- Assessed hold-out and resubstitution estimates after classifier coefficient estimation.
Main Results:
- Resubstitution estimates were consistently optimistically biased, except when too few features were selected.
- Hold-out estimates were always pessimistically biased when feature selection used only design samples.
- When feature selection used all samples, hold-out estimates could be optimistically or pessimistically biased based on sample size, feature count, and data distribution.
- Under specific simulation conditions, hold-out estimates were conservatively biased if the sample-to-feature ratio per class exceeded five.
Conclusions:
- Finite sample size and the method of feature selection (design-only vs. all samples) critically influence classification accuracy and bias in performance estimates.
- Careful consideration of sample size and feature selection strategy is crucial for reliable CAD system development.
- The observed biases highlight the need for robust validation techniques in machine learning for medical diagnosis.

