Related Experiment Video
Updated: Jul 11, 2025

Stepwise Dosing Protocol for Increased Throughput in Label-Free Impedance-Based GPCR Assays
Published on: February 21, 2020
Expanding Training Data for Structure-Based Receptor-Ligand Binding Affinity Regression through Imputation of Missing
Paul G Francoeur1, David R Koes1
1Department of Computational and Systems Biology, University of Pittsburgh, Pittsburgh, Pennsylvania 15260, United States.
This study explores expanding limited molecular property prediction data by imputing binding affinity labels. This machine learning approach shows a small performance improvement for structure-based models despite added noise.
Area of Science:
- Computational chemistry
- Machine learning
- Drug discovery
Background:
- Machine learning models require large datasets for effective training.
- Current structure-based molecular property prediction models face limited training data, particularly for receptor-ligand binding affinity.
- Existing datasets like CrossDocked2020 expanded binding pose classification data but not binding affinity data.
Purpose of the Study:
- To investigate the viability of imputing binding affinity labels to expand training data for structure-based molecular property prediction models.
- To assess the impact of using imputed labels on the performance of binding affinity regression models.
Main Methods:
- Utilized imputed binding affinity labels for molecular complexes lacking experimental data.
- Trained a convolutional neural network using existing binding affinity data from the CrossDocked2020 dataset.
- Evaluated the performance of structure-based models trained with both experimental and imputed binding affinity data.
Main Results:
- Imputing binding affinity labels is a feasible strategy for augmenting training datasets.
- Models trained with imputed labels demonstrated a marginal improvement in binding affinity regression performance.
- The inclusion of imputed labels introduced additional noise into the training data, yet performance still improved.
Conclusions:
- Imputation of binding affinity labels offers a practical method to increase training data for structure-based molecular modeling.
- Despite inherent noise, imputed data can enhance the predictive accuracy of machine learning models for receptor-ligand binding affinity.
- This approach contributes to advancing computational drug discovery by overcoming data limitations.
More Related Videos
08:31Biosensor-based High Throughput Biopanning and Bioinformatics Analysis Strategy for the Global Validation of Drug-protein Interactions
Published on: December 1, 2020
10:29Quantitative Structure-Activity Relationship, Activity Prediction, and Molecular Dynamics of Non-nucleotide Reverse Transcriptase Inhibitors
Published on: May 9, 2025
Related Concept Videos
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
The Equilibrium Binding Constant and Binding Strength
Mechanistic Models: Compartment Models in Individual and Population Analysis
Quantitative Aspects of Drug-Receptor Interaction
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...