Related Experiment Videos
Optimal number of features as a function of sample size for various classification rules
Jianping Hua1, Zixiang Xiong, James Lowey
1Department of Electrical Engineering, Texas A&M University, College Station, TX 77843, USA.
Bioinformatics (Oxford, England)
|December 2, 2004
Summary
Finding the optimal number of features is crucial for classification accuracy, especially with small sample sizes common in gene expression studies. This research uses simulations to determine the best feature count for various classifiers and data distributions.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning
Background:
- Classification error typically decreases with more features, but can increase with too many features, especially in small sample datasets.
- Gene-expression-based phenotype discrimination often involves small sample sizes, exacerbating the challenge of selecting an optimal number of features.
- The issue of finding an optimal number of features is critical for fixed sample sizes and feature-label distributions.
Purpose of the Study:
- To investigate the relationship between feature count, sample size, and classification error.
- To identify the optimal number of features for various classification rules and data distributions.
- To provide a resource for small-sample classification challenges.
Main Methods:
- Employed simulation across diverse feature-label distributions, classification rules, and sample/feature-set sizes.
- Utilized massively parallel computation to find the optimal number of features as a function of sample size.
- Assessed seven classifiers including 3-nearest-neighbor, SVMs, and linear discriminant analysis, with three Gaussian-based models.
- Incorporated real breast cancer patient data and assumed blocked covariance matrices for correlated gene subsets.
Main Results:
- Generated numerous error surfaces illustrating classification error as a function of feature number and sample size.
- These error surfaces are available on a companion website for researchers.
- The study provides insights into classifier performance across different data characteristics and feature set sizes.
Conclusions:
- The study addresses the critical problem of optimal feature selection in small-sample classification scenarios.
- Simulation and computational analysis reveal patterns in classification error related to feature and sample sizes.
- A comprehensive resource of error surfaces is provided to aid researchers in gene-expression-based classification.