Related Experiment Video
Updated: Jun 15, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Estimation of the applicability domain of kernel-based machine learning models for virtual screening.
Nikolas Fechner1, Andreas Jahn, Georg Hinselmann
1Center for Bioinformatics Tübingen (ZBIT), University of Tübingen, Sand 1, 72076 Tübingen, Germany. nikolas.fechner@uni-tuebingen.de.
We developed new methods to determine the applicability domain for kernel-based quantitative structure-activity relationship (QSAR) models. These methods improve virtual screening by identifying reliable predictions and excluding unreliable ones.
Area of Science:
- Computational chemistry
- Cheminformatics
- Machine learning
Background:
- Virtual screening relies on quantitative structure-activity relationship (QSAR) models but struggles with diverse chemical datasets.
- Current applicability domain methods often use vector descriptors, limiting compatibility with structured kernel-based QSAR models.
- Defining the chemical space where QSAR models are reliable is crucial for accurate predictions.
Purpose of the Study:
- To propose and evaluate novel methods for estimating the applicability domain of kernel-based QSAR models.
- To address the limitations of existing applicability domain approaches for structured kernel methods.
Main Methods:
- Developed three distinct approaches for applicability domain estimation tailored for kernel-based QSAR.
- Utilized support vector regression with three different structured kernels on three virtual screening tasks.
- Quantitatively assessed applicability using a score for each compound in the screening dataset.
Main Results:
- The proposed methods successfully differentiated reliable prediction regions within the chemical space.
- Evaluating model performance on subsets defined by applicability scores showed clear separation.
- Omitting compounds with low applicability scores significantly improved virtual screening performance.
Conclusions:
- The developed applicability domain formulations effectively identify compounds unsuitable for reliable QSAR prediction.
- Excluding unreliable predictions enhances virtual screening efficiency by reducing search space.
- The identified compounds would likely not be predicted as active by the model anyway, thus not hindering discovery.
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Estimation of k and VD of Aminoglycosides
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...