Tapping on the Black Box: How Is the Scoring Power of a Machine-Learning Scoring Function Dependent on the Training

Minyi Su1,2, Guoqin Feng1,2, Zhihai Liu1

  • 1State Key Laboratory of Bioorganic and Natural Products Chemistry, Center for Excellence in Molecular Synthesis, Shanghai Institute of Organic Chemistry, Chinese Academy of Sciences, 345 Lingling Road, Shanghai 200032, People's Republic of China.

Summary

Machine-learning scoring functions for protein-ligand interactions show performance dependent on training data. Random Forest models learned best, unlike conventional functions, highlighting the need to consider training-test set similarity for accurate evaluation.

Related Concept Videos

Survival Tree01:19

Survival Tree

Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
337
Outliers and Influential Points01:08

Outliers and Influential Points

An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
5.7K
Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Weighted Mean00:57

Weighted Mean

While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
6.1K
Degrees of Freedom01:02

Degrees of Freedom

The degree of freedom for a particular statistical calculation is the number of values that are free to vary. Thus, the minimum number of independent numbers can specify a particular statistic. The degrees of freedom differ greatly depending on known and uncalculated statistical components.
For example, suppose there are three unknown numbers whose mean is 10; although we can freely assign values to the first and second numbers, the value of the last number can not be arbitrarily assigned.
6.4K
Degrees of Freedom01:02

Degrees of Freedom

The degree of freedom for a particular statistical calculation is the number of values that are free to vary. As a result, the minimum number of independent numbers can specify a particular statistic. The degrees of freedom differ greatly depending on known and uncalculated statistical components.
For example, suppose there are three unknown numbers whose mean is 10; although we can freely assign values to the first and second numbers, the value of the last number can not be arbitrarily...
8.8K