Beware of machine learning-based scoring functions-on the danger of developing black boxes
Joffrey Gabel1, Jérémy Desaphy, Didier Rognan
1Laboratoire d'Innovation Thérapeutique, UMR 7200 CNRS-Université de Strasbourg , 74 route du Rhin, F-67400 Illkirch, France.
Journal of Chemical Information and Modeling
|September 11, 2014
Summary
Machine learning scoring functions trained on simple protein-ligand distance counts show poor virtual screening performance. These models are insensitive to docking pose accuracy, unlike empirical methods, necessitating careful benchmarking for reliable drug discovery.
Area of Science:
- Computational chemistry
- Cheminformatics
- Machine learning in drug discovery
Background:
- Machine learning (ML) algorithms using protein-ligand descriptors show promise for predicting binding constants.
- Recent studies suggest ML approaches outperform traditional empirical scoring functions.
Purpose of the Study:
- To validate the claimed superiority of ML-based scoring functions over empirical methods.
- To assess the performance of ML scoring functions in virtual screening enrichment.
- To investigate the sensitivity of ML scoring functions to ligand pose accuracy.
Main Methods:
- Trained Random Forest and Support Vector Machine models on protein-ligand distance counts from the PDBBind dataset.
- Evaluated scoring functions using 10 DUD-E datasets for virtual screening enrichment.
- Systematically varied ligand poses from X-ray coordinates to test pose sensitivity.
Main Results:
- ML scoring functions accurately predicted binding constants but failed to enrich virtual screening hit lists.
- Empirical Surflex-Dock showed sensitivity to docking pose quality, while ML models were insensitive.
- ML models, trained on basic distance counts, demonstrated insensitivity to ligand pose variations up to 10 Å RMSD.
Conclusions:
- While ML is valuable for scoring function design, using simple protein-ligand element-element distance counts requires caution.
- Proposed two essential benchmarking tests for novel scoring functions: pose sensitivity and virtual screening enrichment capability.
- Emphasized the need for meaningful application of ML descriptors to avoid developing ineffective scoring functions.
Related Concept Videos
Weighted Mean
5.4K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
5.4K
Introduction to z Scores
1.5K
A z score (or standardized value) is measured in units of the standard deviation. It indicates how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
z scores...
1.5K
Introduction to z Scores
8.3K
A z score (or standardized value) is measured in units of the standard deviation. It tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a zero z score. It is important to note that the mean of the z scores is zero, and the standard deviation is one.
z scores...
z scores...
8.3K
Aggregates Classification
991
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
991
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
435
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
435
