Related Experiment Video
Updated: Mar 17, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Selecting Relevant Descriptors for Classification by Bayesian Estimates: A Comparison with Decision Trees and Support
Miriam Carbon-Mangels1, Michael C Hutter2
1Section of Biostatistics, Paul-Ehrlich-Institut, Federal Institute for Vaccines and Biomedicines, Paul-Ehrlich-Straße 51-59, 63225 Langen, Germany.
Bayesian estimates effectively identify relevant molecular descriptors for classification, outperforming other methods on unbalanced datasets. This approach mitigates overfitting by focusing on descriptors with Gaussian-like distributions and high discriminative power.
Area of Science:
- Computational chemistry
- Machine learning in drug discovery
Background:
- Classification algorithms face overfitting due to high dimensionality.
- Identifying relevant descriptors is crucial for reducing complexity and improving model performance.
Purpose of the Study:
- To apply Bayesian estimates for modeling descriptor probability distributions in binary classification.
- To assess the discriminative power of descriptors using Kullback-Leibler divergence.
- To compare Bayesian estimates with LASSO, decision trees, and support vector machines on unbalanced datasets.
Main Methods:
- Utilized Bayesian estimates to model descriptor value distributions for binary classification.
- Employed n-fold cross-validation to evaluate classifier performance.
- Calculated symmetric Kullback-Leibler divergence to measure discriminative power.
- Compared results with LASSO feature selection, decision trees, and support vector machines.
Main Results:
- Relevant descriptors exhibit Gaussian-like distributions and large Kullback-Leibler divergences.
- Bayesian estimates demonstrate robustness against unbalanced data, unlike decision trees and support vector machines.
- Identified descriptors enabling linear class separation.
Conclusions:
- Bayesian estimates provide a robust method for feature selection in classification tasks, particularly with imbalanced data.
- The approach effectively identifies descriptors that facilitate linear separation, reducing the risk of overfitting.
- This strategy aids in understanding the underlying chemical space relevant for classification.
Related Concept Videos
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Methods of Classification and Identification
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...