Related Experiment Videos
Structural and Sequence Similarity Makes a Significant Impact on Machine-Learning-Based Scoring Functions for
Yang Li1,2, Jianyi Yang2
1College of Life Sciences, Nankai University , Tianjin 300071, China.
Journal of Chemical Information and Modeling
|March 31, 2017
Summary
Machine-learning scoring functions for protein-ligand binding affinity are highly accurate due to training data similarity. Removing similar proteins significantly reduces their performance, unlike conventional functions.
Area of Science:
- Computational chemistry
- Structural biology
- Machine learning
Background:
- Machine-learning (ML) scoring functions show remarkable improvements in predicting protein-ligand binding affinity.
- These ML methods, like RF-Score, achieve high correlation coefficients on benchmark datasets, outperforming conventional scoring functions.
- The underlying reasons for the superior performance of ML-based methods remain unclear.
Purpose of the Study:
- To investigate the impact of protein structural and sequence similarity on the performance of ML-based scoring functions.
- To determine whether the high performance of ML methods is an artifact of training data similarity.
Main Methods:
- Systematically controlled structural and sequence similarity between training and test proteins using the PDBbind benchmark.
- Utilized structure and sequence alignment to identify and remove highly similar proteins from training sets.
- Compared the performance of ML-based methods and conventional scoring functions (e.g., X-Score) under varying similarity conditions.
Main Results:
- Protein structural and sequence similarity significantly impacts the performance of ML-based scoring functions.
- When highly similar training proteins were removed, ML methods no longer outperformed conventional scoring functions.
- Conventional scoring functions, such as X-Score, demonstrated relatively stable performance regardless of training data similarity.
Conclusions:
- The high predictive power of current ML-based scoring functions may be largely attributed to the similarity between training and test proteins.
- This highlights the importance of careful dataset curation and validation when developing and benchmarking ML methods for protein-ligand binding affinity prediction.
- Conventional scoring functions exhibit more robust performance across different training datasets.