Related Experiment Video
Updated: Aug 10, 2025

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
Improving classification of correct and incorrect protein-protein docking models by augmenting the training set
Didier Barradas-Bautista1, Ali Almajed2, Romina Oliva3
1Kaust Visualization Lab, Core Lab Division, King Abdullah University of Science and Technology (KAUST), Thuwal 23955-6900, Saudi Arabia.
A new weak supervision method, hAIkal, enhances protein-protein docking by augmenting data for machine learning. This improves the accuracy of predicting protein interaction poses, surpassing current state-of-the-art methods.
Area of Science:
- Computational biology
- Structural bioinformatics
- Machine learning in bioinformatics
Background:
- Protein-protein interactions are crucial for biological processes.
- Experimental determination of 3D structures is time-consuming and costly.
- Computational protein-protein docking generates numerous poses, creating imbalanced datasets unsuitable for machine learning.
Purpose of the Study:
- To address the challenge of imbalanced datasets in protein-protein docking.
- To develop a data augmentation method for improving machine learning model training.
- To enhance the accuracy of protein-protein docking pose prediction.
Main Methods:
- Developed a weak supervision-based data augmentation technique named hAIkal.
- Utilized hAIkal to increase the volume of labeled training data.
- Trained and evaluated multiple machine learning classifiers on the augmented dataset.
Main Results:
- The hAIkal method significantly increased the labeled training data for docking classifiers.
- The best-trained classifier achieved 81% accuracy and a 0.51 Matthews' correlation coefficient on the test set.
- The developed classifier outperformed existing state-of-the-art scoring functions in protein-protein docking.
Conclusions:
- Weak supervision and data augmentation, as implemented in hAIkal, are effective for improving protein-protein docking.
- The enhanced classifiers provide a more accurate prediction of protein interaction poses.
- This approach offers a viable computational solution to complement experimental structure determination.
More Related Videos
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
10:21Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Related Concept Videos
Protein-protein Interfaces
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Organization
The primary structure of a protein is its amino acid sequence....