Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Conserved Binding Sites01:49

Conserved Binding Sites

4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Ligand Binding Sites02:40

Ligand Binding Sites

12.9K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
12.9K
The Equilibrium Binding Constant and Binding Strength02:18

The Equilibrium Binding Constant and Binding Strength

12.9K
The equilibrium binding constant (Kb) quantifies the strength of a protein-ligand interaction. Kb can be calculated as follows when the reaction is at equilibrium:
12.9K
Protein-protein Interfaces02:04

Protein-protein Interfaces

12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Protein-Drug Binding: Determination Methods01:22

Protein-Drug Binding: Determination Methods

210
Determining protein-drug binding can be achieved through indirect and direct methods, each providing valuable insights into the interaction between proteins and drugs.
Indirect methods involve isolating the bound drug from its free form in biological samples such as blood, serum, or plasma. These techniques aim to measure the percentage of drugs bound to proteins. Equilibrium dialysis is a commonly used method where the free drug concentration at equilibrium is measured by separating the bound...
210
Protein Organization01:24

Protein Organization

6.5K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
6.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Large-Scale Collaborative Assessment of Binding Free Energy Calculations for Drug Discovery Using OpenFE.

Journal of chemical information and modeling·2026
Same author

Discovery of 5‑Azaindole Inhibitors of O‑GlcNAcase for the Treatment of Alzheimer's Disease and Related Tauopathies.

ACS medicinal chemistry letters·2026
Same author

Frags2Drugs: A Novel In Silico Fragment-Based Approach to the Discovery of Kinase Inhibitors.

Pharmaceuticals (Basel, Switzerland)·2026
Same author

REINFORCE-ING Chemical Language Models for Drug Discovery.

Journal of chemical information and modeling·2025
Same author

C2PO: an ML-powered optimizer of the membrane permeability of cyclic peptides through chemical modification.

Journal of cheminformatics·2025
Same author

MolAgent: Biomolecular Property Estimation in the Agentic Era.

Journal of chemical information and modeling·2025

Related Experiment Video

Updated: Jul 10, 2025

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
10:21

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA

Published on: February 23, 2024

2.6K

The Impact of Data on Structure-Based Binding Affinity Predictions Using Deep Neural Networks.

Pierre-Yves Libouban1, Samia Aci-Sèche1, Jose Carlos Gómez-Tamayo2

  • 1Institute of Organic and Analytical Chemistry (ICOA), UMR7311, Université d'Orléans, CNRS, Pôle de Chimie rue de Chartres, 45067 Orléans, CEDEX 2, France.

International Journal of Molecular Sciences
|November 25, 2023
PubMed
Summary

This study examines how the quality and quantity of training data influence the accuracy of artificial intelligence models designed to predict how strongly drugs bind to target proteins. The researchers found that the size of the protein binding site is a key factor, and that using larger, more diverse datasets is more beneficial than focusing solely on high-quality data. They also highlight that current testing methods may be biased, suggesting a need for more rigorous benchmarking to truly evaluate these tools.

Keywords:
binding affinitiesdeep learningprotein–liganddrug discoveryartificial intelligenceprotein-ligand interactioncomputational modeling

Frequently Asked Questions

More Related Videos

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
08:49

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis

Published on: June 20, 2025

196
Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
06:50

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions

Published on: January 26, 2024

1.9K

Related Experiment Videos

Last Updated: Jul 10, 2025

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
10:21

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA

Published on: February 23, 2024

2.6K
Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
08:49

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis

Published on: June 20, 2025

196
Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
06:50

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions

Published on: January 26, 2024

1.9K

Area of Science:

  • Computational chemistry within drug discovery research
  • Deep learning applications in binding affinity prediction

Background:

No prior work had fully resolved why performance plateaus exist in computational drug design models. It was already known that artificial intelligence algorithms are frequently utilized to estimate protein-ligand interaction strengths. That uncertainty drove researchers to question if current successes reflect genuine progress or merely reliance on flawed information. Prior research has shown that neural network architectures and system representations are often prioritized over data quality. This gap motivated a closer look at the underlying training sets used for model development. Many existing studies assume that high-quality data alone will yield superior predictive power. However, the influence of specific input parameters on these outcomes remains poorly understood. This investigation addresses the potential for biased, easily predictable data to skew results in the field.

Purpose Of The Study:

The aim of this work is to investigate how specific input data parameters influence the performance of binding affinity prediction models. Researchers sought to determine if current predictive plateaus result from inherent model limitations or biased training information. They addressed the uncertainty surrounding whether high-quality data is always superior to larger, more diverse datasets. This study explores the impact of binding pocket size on the statistical accuracy of neural networks. The team intended to clarify whether current evaluation benchmarks provide an accurate representation of model capabilities. They motivated this analysis by questioning the optimistic performance reports common in existing literature. The researchers aimed to provide a clearer understanding of how data selection shapes the decision-making processes of these algorithms. This effort was driven by the need to establish more reliable protocols for evaluating computational drug discovery tools.

Main Methods:

The review approach involved systematically testing various input parameters to determine their influence on predictive accuracy. Researchers constructed multiple statistical models to observe how different data configurations altered output results. They evaluated the impact of binding pocket dimensions on the overall efficacy of the computational systems. The team compared models trained on expansive datasets against those restricted to high-quality subsets. They performed rigorous benchmarking to identify potential biases within common testing protocols. The study utilized deep learning frameworks to process complex protein-ligand interactions. They analyzed how different training strategies affected the decision-making processes of these algorithms. The team synthesized these observations to provide a clear picture of how data characteristics shape model performance.

Main Results:

Key findings from the literature indicate that the dimensions of the binding pocket represent a primary factor influencing predictive success. The authors demonstrate that training models with maximum data volume yields better results than limiting inputs to high-quality samples. They confirm that standard test sets currently in use contain significant biases that skew performance metrics. The research shows that current predictive performance has reached a plateau across many common architectures. They identify that relying on easily predictable data leads to overly optimistic outcomes in model evaluation. The study reveals that diverse data sources are more effective than refined subsets for improving model generalization. They report that existing evaluation methods fail to capture the true decision-making capabilities of these systems. The findings suggest that current benchmarks are insufficient for accurately comparing different computational approaches.

Conclusions:

The authors propose that multiple benchmarking strategies are necessary to properly assess model decision-making. Their findings suggest that the dimensions of the binding pocket significantly dictate predictive success. They confirm that existing test sets often contain inherent biases that misrepresent true model capabilities. The researchers conclude that prioritizing data volume over strict quality filters improves overall performance. This synthesis implies that current evaluation metrics may provide an overly optimistic view of model utility. They suggest that future efforts should focus on diverse data rather than just refined subsets. The evidence indicates that understanding the limitations of input data is vital for progress. These implications highlight the necessity of rigorous validation to ensure reliable drug discovery outcomes.

The researchers propose that the size of the binding pocket acts as a primary determinant for predictive success. Larger pockets provide more structural information, which helps the neural network distinguish between different ligand interactions compared to smaller, less informative sites.

The authors utilized deep learning architectures to process protein-ligand complexes. These models were trained on varying amounts of data to evaluate how input volume versus quality affects the final predictive accuracy of the system.

A diverse, large-scale dataset is necessary to overcome the performance plateaus observed in current models. The authors argue that restricting training to high-quality subsets limits the model's ability to generalize, unlike using expansive, varied data sources.

The researchers used these sets to identify systemic biases in how models are currently evaluated. They found that standard benchmarks often produce overly optimistic results, unlike more rigorous, multi-faceted testing protocols that reveal true model limitations.

The study measured the predictive accuracy of binding affinity models across different training conditions. They observed that models trained on larger, more inclusive datasets outperformed those limited to high-quality samples, demonstrating a clear performance gap.

The authors imply that current evaluation methods are insufficient for understanding model decision-making. They suggest that moving beyond standard benchmarks is necessary to accurately compare different architectures, rather than relying on existing, potentially flawed metrics.