Related Experiment Video
Updated: May 23, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Bootstrap-based K-means feature selection strategy for fuzzy regression functions
Aylin Ucan1,2, Dogan Yildiz1, Nihat Tak3
1Department of Statistics, Yildiz Technical University, Istanbul, 34220, Turkey.
This study introduces a bootstrap-stabilized clustering method for robust feature selection in complex datasets. The approach enhances model generalizability and offers competitive performance compared to existing techniques.
Area of Science:
- Data Science
- Machine Learning
- Statistics
Background:
- Traditional feature selection methods struggle with complex datasets.
- Need for stable and robust feature selection alternatives is growing.
- Existing methods can be sensitive to random initialization and lack generalizability.
Purpose of the Study:
- Propose a novel bootstrap-stabilized feature selection method.
- Combine stable feature selection with Fuzzy Regression Function (FRF) modeling.
- Evaluate the proposed method's effectiveness on real-world datasets.
Main Methods:
- Utilized k-means clustering for grouping features across multiple resamples.
- Identified stable representative features by assigning them to nearest cluster centers.
- Integrated selected features into a Fuzzy Regression Function (FRF) model.
- Determined the number of feature clusters using grid-search, Silhouette Index (SI), and Davies Bouldin Index (DBI).
Main Results:
- The proposed bootstrap-stabilized clustering approach demonstrated superior or competitive performance.
- Effectiveness was validated across ten diverse real-world datasets.
- The hybrid framework combines interpretable feature selection with flexible fuzzy regression.
Conclusions:
- The bootstrap-stabilized clustering method is effective for feature selection in complex data.
- The approach offers improved stability and generalizability over traditional methods.
- This hybrid framework is suitable for complex data environments requiring robust feature selection.
Related Concept Videos
Regression Toward the Mean
Quantifying and Rejecting Outliers: The Grubbs Test
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Expected Frequencies in Goodness-of-Fit Tests
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
