Related Experiment Video
Updated: Jan 9, 2026

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Can AI-Predicted Complexes Teach Machine Learning to Compute Drug Binding Affinity?
Wei-Tse Hsu1, Savva Grevtsev2, Anna M Herz3
1Structural Bioinformatics and Computational Biochemistry, Department of Biochemistry, University of Oxford, South Parks Road, Oxford, OX1 3QU, U.K.
Co-folding models can augment data for machine learning-based scoring functions (MLSFs). Performance gains from synthetic data depend on structural quality, guiding future data augmentation strategies.
Area of Science:
- Computational chemistry
- Structural biology
- Machine learning
Background:
- Machine learning-based scoring functions (MLSFs) are crucial for predicting binding affinity.
- Synthetic data augmentation is a promising approach to improve MLSF performance.
- Co-folding models offer a potential source of synthetic structural data.
Purpose of the Study:
- To evaluate the feasibility of using co-folding models for synthetic data augmentation in MLSF training.
- To determine the impact of augmented data quality on MLSF performance.
- To develop methods for identifying high-quality co-folding predictions for data augmentation.
Main Methods:
- Utilized co-folding models to generate synthetic protein complex structures.
- Trained MLSFs using both experimental and co-folding-derived data.
- Developed and applied heuristics to assess the structural quality of co-folding predictions.
- Compared MLSF performance using different data augmentation strategies.
Main Results:
- Performance gains from co-folding data augmentation are highly dependent on the structural quality of the predictions.
- Established simple heuristics can effectively identify high-quality co-folding predictions without requiring experimental structures.
- Co-folding predictions, when filtered by quality, can successfully substitute for experimental structures in MLSF training.
Conclusions:
- Co-folding models are a feasible, albeit quality-dependent, source for synthetic data augmentation in MLSF training.
- Heuristics for quality control are essential for successful data augmentation using co-folding models.
- This work provides a framework for leveraging co-folding models to enhance binding affinity prediction through data augmentation.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein-protein Interfaces
The Equilibrium Binding Constant and Binding Strength
The Equilibrium Binding Constant and Binding Strength
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-Drug Binding: Determination Methods
Indirect methods involve isolating the bound drug from its free form in biological samples such as blood, serum, or plasma. These techniques aim to measure the percentage of drugs bound to proteins. Equilibrium dialysis is a commonly used method where the free drug concentration at equilibrium is measured by separating the bound...

