Related Experiment Video
Updated: Jan 15, 2026

Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
Training bias and sequence alignments shape protein-peptide docking by AlphaFold and related methods
Lindsey Guan1, Amy E Keating2,3,4
1Graduate Program in Computational and Systems Biology, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA.
Deep learning models like AlphaFold3 accurately predict protein-peptide structures but show bias towards known data. Unpaired sequence alignments, not coevolution, improve predictions, highlighting the need for diverse training data.
Area of Science:
- Structural biology
- Computational biology
- Bioinformatics
Background:
- Protein-peptide interactions are crucial for biological processes.
- Accurate structural models are vital for understanding function and designing inhibitors.
- Computational models like AlphaFold3 show promise for predicting these structures.
Purpose of the Study:
- To analyze the performance of four protein-peptide structure prediction models: AlphaFold2-Multimer, AlphaFold3, Boltz-1, and Chai-1.
- To investigate how these models utilize multiple sequence alignments (MSAs) for predictions.
- To identify limitations and areas for improvement in deep learning-based peptide docking.
Main Methods:
- Evaluation of four protein-peptide structure prediction models using a dataset of experimentally resolved structures.
- Analysis of model performance regarding generalization to novel protein-peptide complexes.
- Investigation of the contribution of protein and peptide unpaired and paired MSAs to prediction accuracy.
Main Results:
- Models demonstrated high accuracy but exhibited bias towards previously observed structures, limiting generalization to novel complexes.
- Shallow or poor-quality MSAs for peptides were noted.
- Weak evidence for the use of coevolutionary information from paired MSAs was found.
- Both protein and peptide unpaired MSAs significantly contributed to prediction accuracy.
Conclusions:
- Deep learning models show significant promise for protein-peptide docking.
- Model performance is influenced by the quality and diversity of training data, particularly interface geometries.
- Future improvements require addressing model biases and enhancing the representation of novel interactions in training datasets.
Related Concept Videos
Protein Folding
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Protein Folding
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein Organization
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...

