Related Experiment Video
Updated: Sep 22, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Robust Random Forest-Based All-Relevant Feature Ranks for Trustworthy AI
Bastian Pfeifer1, Andreas Holzinger1,2, Michael G Schimek1
1Institute for Medical Informatics Statistics and Documentations, Medical University of Graz, Austria.
Abstract:
Feature selection is a fundamental challenge in machine learning. For instance in bioinformatics, it is essential when one wishes to detect biomarkers. Tree-based methods are predominantly used for this purpose. In this paper, we study the stability of the feature selection methods BORUTA, VITA, and RRF (regularized random forest). In particular, we investigate the feature ranking instability of the associated stochastic algorithms. For stabilization of the feature ranks, we propose to compute consensus values from multiple feature selection runs, applying rank aggregation techniques. Our results show that these consolidated features are more accurate and robust, which helps to make practical machine learning applications more trustworthy.
More Related Videos
Related Concept Videos
Distribution Reliability and Automation
Confidence Coefficient
Random and Systematic Errors
Randomized Experiments
Simple randomization
Simple...
Ranks
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...

