Related Experiment Videos
How wrong can we get? A review of machine learning approaches and error bars
Anton Schwaighofer1, Timon Schroeter, Sebastian Mika
1Technische Universität Berlin, Department of Computer Science, Franklinstrasse 28/29, D-10587 Berlin, Germany.
Combinatorial Chemistry & High Throughput Screening
|June 13, 2009
Summary
This study explores three nonlinear machine learning methods—support vector regression, Gaussian process models, and decision trees—for ligand-based virtual screening. It details obtaining confidence estimates and practical model evaluation for drug discovery applications.
Area of Science:
- Computational chemistry and cheminformatics
- Machine learning applications in drug discovery
Background:
- Ligand-based virtual screening is crucial for identifying potential drug candidates.
- Numerous machine learning methods exist, but their application in virtual screening requires careful consideration.
Purpose of the Study:
- To introduce and compare three nonlinear machine learning methods for ligand-based virtual screening: support vector regression, Gaussian process models, and decision trees.
- To explain how confidence estimates (error bars) can be derived from these models.
- To discuss essential aspects of model building, evaluation, and practical application.
Main Methods:
- Focus on three nonlinear machine learning techniques: support vector regression, Gaussian process models, and decision trees.
- Detailed explanations of each method's principles and intuitive understanding.
- Methodologies for model selection, performance evaluation, and verification of error bar quality.
Main Results:
- Provides an introduction to support vector regression, Gaussian process models, and decision trees for virtual screening.
- Explains the generation and verification of confidence estimates (error bars) for these models.
- Discusses practical considerations for implementing these machine learning methods.
Conclusions:
- The study offers a practical guide to applying specific nonlinear machine learning methods in ligand-based virtual screening.
- Highlights the importance of confidence estimates for robust model evaluation.
- Aims to facilitate the effective use of these computational tools in drug discovery.
Related Concept Videos
Systematic Error: Methodological and Sampling Errors
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
Random and Systematic Errors
Scientists always try their best to record measurements with the utmost accuracy and precision. However, sometimes errors do occur. These errors can be random or systematic. Random errors are observed due to the inconsistency or fluctuation in the measurement process, or variations in the quantity itself that is being measured. Such errors fluctuate from being greater than or less than the true value in repeated measurements. Consider a scientist measuring the length of an earthworm using a...
Accuracy and Errors in Hypothesis Testing
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Random Error
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
Detection of Gross Error: The Q Test
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...