Related Experiment Video
Updated: Dec 28, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Tapping on the Black Box: How Is the Scoring Power of a Machine-Learning Scoring Function Dependent on the Training
Minyi Su1,2, Guoqin Feng1,2, Zhihai Liu1
1State Key Laboratory of Bioorganic and Natural Products Chemistry, Center for Excellence in Molecular Synthesis, Shanghai Institute of Organic Chemistry, Chinese Academy of Sciences, 345 Lingling Road, Shanghai 200032, People's Republic of China.
Machine-learning scoring functions for protein-ligand interactions show performance dependent on training data. Random Forest models learned best, unlike conventional functions, highlighting the need to consider training-test set similarity for accurate evaluation.
Area of Science:
- Computational chemistry
- Drug discovery
- Machine learning in cheminformatics
Background:
- Machine-learning (ML) scoring functions are increasingly used for protein-ligand interactions.
- Concerns exist regarding the dependency of ML scoring function performance on training and test set overlap.
- Systematic evaluation is needed to understand ML model generalizability.
Purpose of the Study:
- To systematically assess the performance of six ML algorithms for protein-ligand scoring.
- To investigate the impact of training set size and similarity on ML scoring function accuracy.
- To compare ML scoring functions against conventional methods under varying data conditions.
Main Methods:
- Six ML algorithms (BRR, DT, KNN, MLP, L-SVR, RF) were trained on subsets of the PDBbind refined set (2016).
- Scoring functions were evaluated on the CASF-2016 test set.
- Training set size and similarity to the test set were systematically varied.
- Performance was compared to conventional scoring functions (ChemScore, ASP, X-Score).
Main Results:
- ML model performance, particularly Random Forest, was dependent on training set size and similarity.
- Conventional scoring functions showed consistent performance regardless of training set variations.
- Random Forest exhibited the best learning capability among the tested ML algorithms.
- "Soft overlap" between training and test sets is crucial for evaluating ML scoring functions.
Conclusions:
- The performance of ML scoring functions is sensitive to training data characteristics (size, similarity).
- Careful consideration of training-test set overlap is essential for robust evaluation of ML scoring functions.
- Standardized datasets accounting for similarity are proposed for future benchmarking.
- Random Forest demonstrates strong potential for developing accurate protein-ligand scoring functions.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Outliers and Influential Points
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Degrees of Freedom
For example, suppose there are three unknown numbers whose mean is 10; although we can freely assign values to the first and second numbers, the value of the last number can not be arbitrarily assigned.
Degrees of Freedom
For example, suppose there are three unknown numbers whose mean is 10; although we can freely assign values to the first and second numbers, the value of the last number can not be arbitrarily...
