Related Experiment Video
Updated: Aug 8, 2025

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-Generated Protein-Ligand
Francesco Pellicani1, Diego Dal Ben2, Andrea Perali3
1Physics Division, School of Science and Technology, University of Camerino, I-62032 Camerino, MC, Italy.
Machine learning shows promise for drug discovery scoring functions. This study found artificial neural networks perform similarly on experimental and generated data, highlighting potential for per-target models.
Area of Science:
- Computational chemistry
- Machine learning in drug discovery
- Bioinformatics
Background:
- Machine learning (ML) is increasingly used for developing accurate scoring functions in computational drug discovery.
- Concerns exist regarding over-optimistic ML results due to database correlations in training and testing sets.
- Investigating ML performance with diverse datasets is crucial for reliable drug discovery tools.
Purpose of the Study:
- To evaluate the performance of an artificial neural network (ANN) for predicting protein-ligand binding affinity.
- To compare ANN performance using experimental versus computer-generated protein-ligand structures.
- To explore the potential of per-target scoring functions trained on computer-generated data.
Main Methods:
- An artificial neural network was trained and tested for binding affinity prediction.
- Two distinct datasets were used: experimental protein-ligand structures and computer-generated structures.
- Performance was assessed using both random (horizontal) and target-specific (vertical) cross-validation tests.
Main Results:
- The ANN demonstrated similar performance regardless of whether experimental or computer-generated structures were used for training and testing.
- A significant performance decrease was observed in vertical tests (unseen target proteins) compared to horizontal tests.
- Training ANNs on computer-generated databases enabled the development of effective per-target scoring functions.
Conclusions:
- Computer-generated structural databases can be effectively utilized for training machine learning models in drug discovery.
- Per-target scoring functions, trained on specialized datasets, show promising results for specific protein targets.
- This approach offers a viable strategy to mitigate database correlation issues and improve the reliability of computational drug discovery.
Related Concept Videos
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Drug Discovery: Overview
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
The Equilibrium Binding Constant and Binding Strength

