Related Experiment Video
Updated: Jun 24, 2026

05:21
Computerized Adaptive Testing System of Functional Assessment of Stroke
Published on: January 7, 2019
Comparative assessment of scoring functions on a diverse test set
1State Key Laboratory of Bioorganic Chemistry, Shanghai Institute of Organic Chemistry, Chinese Academy of Sciences, Shanghai, P. R. China.
Journal of Chemical Information and Modeling
|April 11, 2009
Summary
This study benchmarks 16 scoring functions for protein-ligand binding in drug design. No single function excels across all evaluation metrics, emphasizing the need for careful selection based on specific needs.
Area of Science:
- Computational chemistry and cheminformatics
- Structural biology and drug discovery
- Bioinformatics and computational drug design
Background:
- Scoring functions are crucial for evaluating protein-ligand interactions in structure-based drug design.
- Numerous scoring functions exist, implemented in commercial software or developed by academic groups.
- A comprehensive benchmark is needed to assess their performance.
Purpose of the Study:
- To comparatively assess the performance of 16 popular scoring functions.
- To evaluate functions based on "docking power", "ranking power", and "scoring power".
- To provide an updated benchmark for scoring function evaluation in drug design.
Main Methods:
- Selected 195 diverse protein-ligand complexes from the PDBbind database.
- Evaluated 16 scoring functions independently of molecular docking or virtual screening.
- Assessed "docking power" using a root-mean-square deviation cutoff of 2.0 Å.
- Analyzed "ranking power" and "scoring power" using experimental binding constants.
Main Results:
- Six scoring functions achieved >70% success rate for "docking power"; consensus schemes improved this to >80%.
- X-Score, DrugScore(CSD), DS::PLP, and SYBYL::ChemScore were top performers for "ranking" and "scoring" power.
- Correlation coefficients between computed and experimental binding constants ranged from 0.545 to 0.644.
- Performance varied across different protein-ligand complex types and evaluation aspects.
Conclusions:
- No single scoring function consistently outperforms others across all evaluated aspects.
- The choice of scoring function should be tailored to the specific application and purpose.
- This study provides valuable insights for selecting appropriate scoring functions in drug design.
Related Concept Videos
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Comparison Tests
An infinite series composed of positive terms may either approach a finite value or increase without bound. Determining which outcome occurs is a central task in calculus, and comparison tests provide structured methods for making this determination. Rather than evaluating a series directly, these tests relate it to another series whose behavior is already known, allowing conclusions to be drawn through logical comparison.The direct comparison test applies to series with positive terms. If each...
Comparing Experimental Results: Student's t-Test
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
Bonferroni Test
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
Goodness-of-Fit Test
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
Test for Homogeneity
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can be stated as...
