Related Experiment Video
Updated: Jun 9, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
LLpowershap: logistic loss-based automated Shapley values feature selection method.
Iqbal Madakkatel1,2, Elina Hyppönen3,4
1Australian Centre for Precision Health, Unit of Clinical and Health Sciences, University of South Australia, Adelaide, 5001, South Australia, Australia. iqbal.madakkatel@unisa.edu.au.
LLpowershap, a novel feature selection method, identifies more informative features with less noise. This method demonstrates superior or comparable predictive performance on real-world data, ranking best among tested methods.
Area of Science:
- Machine Learning
- Bioinformatics
- Data Science
Background:
- Shapley values are crucial for explaining complex machine learning models, aiding in debugging, fairness analysis, and feature selection.
- Existing methods like powershap utilize predictive Shapley values and p-values for feature selection.
- Shapley values ensure fair distribution of feature contributions, considering non-linearities and interactions.
Purpose of the Study:
- Introduce LLpowershap, a novel feature selection method utilizing loss-based Shapley values.
- Enhance the identification of informative features while minimizing noise.
- Improve the calculation of p-values and power for feature selection and iteration estimation.
Main Methods:
- Employs loss-based Shapley values for feature selection.
- Incorporates enhanced p-value and power calculations.
- Evaluated through simulations and benchmarking on real-world datasets.
Main Results:
- LLpowershap identifies a greater number of informative features and fewer noise features compared to existing methods.
- Demonstrates high or comparable predictive performance against other Shapley-based and filter methods.
- Achieved the top mean ranking among seven tested feature selection methods.
Conclusions:
- LLpowershap is an effective wrapper feature selection method.
- Suitable for feature selection in large-scale biomedical datasets and other complex data environments.
- Offers a robust approach for identifying key features in machine learning models.
More Related Videos
04:57Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
Published on: May 16, 2022
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
Related Concept Videos
Quantifying and Rejecting Outliers: The Grubbs Test
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Kaplan-Meier Approach
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...