Related Experiment Video
Updated: Dec 10, 2025

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
Three-Dimensional Convolutional Neural Networks and a Cross-Docked Data Set for Structure-Based Drug Design
Paul G Francoeur1, Tomohide Masuda1, Jocelyn Sunseri1
1Department of Computational and Systems Biology, University of Pittsburgh, Pittsburgh, Pennsylvania 15260, United States.
A new dataset, CrossDocked2020, and machine learning models were developed to improve protein-ligand binding affinity prediction. This advances drug discovery by providing a standard benchmark for evaluating model generalization to new targets.
Area of Science:
- Computational chemistry
- Machine learning in drug discovery
- Structural bioinformatics
Background:
- Predicting protein-ligand binding affinity is crucial for drug discovery.
- Current machine learning models for this task often overestimate generalization ability.
- A lack of standardized datasets hinders reliable model comparison.
Purpose of the Study:
- Introduce the CrossDocked2020 dataset for structure-based machine learning.
- Evaluate grid-based convolutional neural network (CNN) models on this new dataset.
- Establish a standardized benchmark for protein-ligand binding affinity prediction.
Main Methods:
- Generated 22.5 million ligand poses docked into similar binding pockets using the CrossDocked2020 dataset.
- Trained and evaluated various grid-based CNN models, including densely connected CNNs.
- Performed comprehensive analysis on data partitioning, training data quality, and pose sensitivity.
Main Results:
- An ensemble of five densely connected CNNs achieved a root mean squared error of 1.42 and Pearson R of 0.612 for affinity prediction.
- The model demonstrated high performance in binding pose classification (AUC 0.956) and pose selection (68.4% accuracy).
- Demonstrated the impact of data partitioning and the benefit of including lower-quality training data.
Conclusions:
- The CrossDocked2020 dataset provides a standardized benchmark for machine learning in drug discovery.
- The developed CNN models show significant potential for accurate protein-ligand binding affinity prediction.
- This work facilitates community adoption and advances the field of computational drug design.
More Related Videos
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
05:18Quaternary Structure Modeling Through Chemical Cross-Linking Mass Spectrometry: Extending TX-MS Jupyter Reports
Published on: October 20, 2021