Related Experiment Video
Updated: Feb 24, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A Multicriteria Approach to Find Predictive and Sparse Models with Stable Feature Selection for High-Dimensional Data
Andrea Bommert1, Jörg Rahnenführer1, Michel Lang1
1Department of Statistics, TU Dortmund University, 44221 Dortmund, Germany.
Developing accurate predictive models for high-dimensional genetic data requires balancing classification accuracy with stable, minimal feature selection. This study identifies Pearson correlation as a robust stability measure for reliable bioinformatics models.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Genetics
Background:
- High-dimensional data, particularly genetic data, presents challenges for predictive modeling.
- Model interpretability and feature selection stability are crucial for biological insights in bioinformatics.
Purpose of the Study:
- To evaluate and identify optimal criteria for fitting predictive models to high-dimensional datasets.
- To compare various feature selection stability measures and assess their suitability for genetic data analysis.
- To investigate the trade-offs between classification accuracy, feature selection stability, and model complexity.
Main Methods:
- Comparison of multiple feature selection stability measures.
- Assessment of stability measures based on theoretical and empirical properties.
- Utilizing Pareto fronts to analyze multi-objective optimization (accuracy, stability, feature count).
- Focus on measures incorporating corrections for chance and the number of selected features.
Main Results:
- Pearson correlation demonstrated superior theoretical and empirical properties as a stability measure.
- Stability assessment requires measures with corrections for chance or the number of features selected.
- Pareto front analysis revealed that models with stable, reduced feature sets can achieve high predictive accuracy.
Conclusions:
- Pearson correlation is recommended for assessing feature selection stability in high-dimensional genetic data.
- Effective predictive models in bioinformatics must prioritize accuracy, stability, and parsimony.
- The study provides a framework for developing reliable and interpretable models for biological discovery.
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Survival Tree
Building a Survival Tree
Constructing a...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
