Related Experiment Video
Updated: May 14, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Nonparametric IPSS: fast, flexible feature selection with false discovery control
Omar Melikechi1, David B Dunson2, Jeffrey W Miller1
1Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, 02115, United States.
This study introduces Integrated Path Stability Selection (IPSS), a novel feature selection method offering robust false discovery control and improved true positive identification for high-dimensional data. IPSS methods like IPSSGB and IPSSRF demonstrate superior performance in simulations and cancer-related gene discovery.
Area of Science:
- Machine Learning
- Statistical Methods
- Bioinformatics
Background:
- Feature selection is crucial in machine learning and statistics.
- Existing methods often rely on parametric models, lack theoretical false discovery control, or identify limited true positives.
Purpose of the Study:
- Introduce a general, nonparametric feature selection method with finite-sample false discovery control.
- Enhance the identification of true positives while maintaining statistical rigor.
- Provide efficient and accurate tools for high-dimensional data analysis.
Main Methods:
- Utilize Integrated Path Stability Selection (IPSS) applied to arbitrary feature importance scores.
- Develop specific implementations: IPSS Gradient Boosting (IPSSGB) and IPSS Random Forests (IPSSRF).
- Estimate q-values for better suitability in high-dimensional settings compared to P-values.
Main Results:
- IPSSGB and IPSSRF demonstrate accurate false discovery rate control in nonlinear simulations.
- Both methods significantly outperform existing approaches in detecting true positives.
- Achieve high efficiency, running in under 20 seconds for datasets with 500 samples and 5000 features.
- Applied to cancer data, IPSSGB and IPSSRF yield improved predictions with fewer features.
Conclusions:
- IPSS offers a powerful and flexible framework for feature selection in high-dimensional data.
- The developed IPSSGB and IPSSRF methods provide accurate and efficient solutions for biological data analysis.
- These methods advance the field by improving both statistical control and discovery power in feature selection.
More Related Videos
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
09:01A Method for Investigating Age-related Differences in the Functional Connectivity of Cognitive Control Networks Associated with Dimensional Change Card Sort Performance
Published on: May 7, 2014
Related Concept Videos
Introduction to Nonparametric Statistics
One of...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Quantifying and Rejecting Outliers: The Grubbs Test
Fisher's Exact Test
Expected Frequencies in Goodness-of-Fit Tests
Statistical Package for the Social Sciences (SPSS)
SPSS streamlines the process from data preparation to analysis and reporting. It is characterized by its user-friendly interface, which conceals...