Related Experiment Videos
Novel unsupervised feature filtering of biological data
Roy Varshavsky1, Assaf Gottlieb, Michal Linial
1School of Computer Science and Engineering, The Hebrew University of Jerusalem, 91904 Israel. royke@cs.huji.ac.il
Bioinformatics (Oxford, England)
|July 29, 2006
Summary
This study introduces a novel unsupervised feature selection method using SVD-entropy. This approach effectively identifies informative features in large, noisy datasets, outperforming existing techniques for improved data analysis and clustering.
Area of Science:
- Computational Biology
- Data Science
- Bioinformatics
Background:
- Unsupervised feature selection is crucial for large, noisy datasets.
- Existing unsupervised methods are limited.
- Novel methods are needed to identify informative feature subsets.
Purpose of the Study:
- To propose a novel unsupervised feature selection criterion based on SVD-entropy.
- To evaluate the effectiveness of this new method in identifying informative features.
- To compare the proposed method against existing techniques.
Main Methods:
- Developed a novel unsupervised criterion using SVD-entropy.
- Calculated contribution to entropy (CE) on a leave-one-out basis.
- Implemented four variations: simple ranking (SR), forward selection (FS1, FS2), and backward elimination (BE).
Main Results:
- Applied methods to benchmark datasets and evaluated clustering success using Jaccard scores.
- Feature filtering by CE outperformed variance and gene-shaving methods.
- Selected feature sets achieved superior clustering performance compared to using all features in some cases.
- Identified an optimal feature set size, typically a small percentage of total features.
- Selected genes showed significant Gene Ontology (GO) enrichment in relevant cellular processes.
Conclusions:
- The SVD-entropy based feature selection method is effective for unsupervised learning.
- This approach can identify a small subset of highly informative features.
- The selected features are biologically relevant and improve data clustering performance.