Related Experiment Video
Updated: Jun 8, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Leave-cluster-out cross-validation is appropriate for scoring functions derived from diverse protein data sets
Christian Kramer1, Peter Gedeck
1Novartis Institutes for BioMedical Research, Novartis Pharma AG, Forum 1, Novartis Campus, CH-4056 Basel, Switzerland. Christian.Kramer@novartis.com
Accurate validation of scoring functions is crucial for drug discovery. Standard validation methods may overestimate performance when training and testing sets share similar protein families, highlighting the need for rigorous cross-validation techniques.
Area of Science:
- Computational chemistry
- Structural biology
- Drug discovery
Background:
- Large datasets of protein-ligand complexes (e.g., PDBbind) enable new approaches for scoring function development.
- Quantitative Structure-Activity Relationship (QSAR) principles can be applied to scoring functions using machine learning.
Purpose of the Study:
- To evaluate the impact of different validation strategies on scoring function performance.
- To highlight potential overestimation of predictive accuracy with standard validation methods.
Main Methods:
- Development and validation of a scoring function using protein-ligand interaction descriptors.
- Comparison of validation results using the PDBbind core set versus leave-cluster-out cross-validation.
- Analysis of performance metrics (R and R²) under different validation schemes.
Main Results:
- Significant discrepancies in performance metrics (R: 0.77 vs 0.46, R²: 0.59 vs 0.21) were observed based on the validation method.
- Leave-cluster-out cross-validation revealed lower, more realistic performance compared to standard validation.
- Standard validation overestimates scoring function quality when training and validation sets contain related proteins.
Conclusions:
- Rigorous cross-validation is essential for reliable scoring function evaluation.
- Care must be taken to avoid optimistic performance estimates, especially when dealing with homologous proteins in training and validation sets.
- The choice of validation strategy significantly impacts the perceived accuracy of computational scoring functions.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Quantifying and Rejecting Outliers: The Grubbs Test
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Expected Frequencies in Goodness-of-Fit Tests
Protein Folding Quality Check in the RER
Kendall's Coefficient of Concordance
