Related Experiment Video
Updated: Sep 16, 2026

Computational Prediction of Amino Acid Preferences of Potentially Multispecific Peptide-Binding Domains Involved in Protein-Protein Interactions
Published on: January 26, 2024
A curated structural dataset of peptide-protein complexes reveals biases in existing datasets and principles of
Rahma Hamdani1, Javier Delgado1, Luis Serrano1,2,3
1Centre for Genomic Regulation (CRG), The Barcelona Institute for Science and Technology, Barcelona, Spain.
Abstract:
Peptide-protein interactions are fundamental to many biological processes, and peptide design is gaining interest due to its therapeutic potential. This has led to the emergence of various structural databases, such as PepBDB and Peptipedia, that classify peptides as polypeptides with fewer than 50 amino acids. These databases provide valuable starting points for studying peptide recognition by protein partners and are widely used for machine learning applications, docking, and scoring functions. However, under such a length definition for peptides, we have very different cases that are likely to confound the analysis of peptide binding, including miniproteins, peptide-peptide complexes, intramolecular peptide disulfide bonds, intermolecular disulfide bridges, and proteins undergoing internal cleavage, such as serpins, as well as non-natural amino acids and covalently bound cofactors. Here, we present a rigorous classification of peptide-protein complexes to generate datasets suitable for comparative energetic analysis with a focus on peptides that are unstructured in the absence of their target protein. The analysis of this dataset shows that peptide binding is typically driven by a small number of hotspot residues mainly enriched in aromatic and bulky hydrophobic side chains. Their number of hotspots and their spatial organization depend on peptide length, secondary structure, and covalent constraints. Short peptides rely on central anchor regions, whereas longer peptides distribute hotspots more broadly, with helices showing periodic spacing and β-strands relying more on backbone-mediated stabilization. Disulfide bonds further decrease the number of hotspots per peptide length by either pre-organizing the peptide or acting as covalent anchors. This work provides a curated resource and general principles for peptide recognition. It highlights the importance of structurally classifying peptide-protein complexes to avoid bias in downstream computational and machine-learning applications.
Related Concept Videos
Protein-protein Interfaces
Protein-Protein Interfaces
Protein Organization
The primary structure of a protein is its amino acid sequence.
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...

