High fitness paths can connect proteins with low sequence overlap
Pranav Kantroo1,2, Günter P Wagner3,4,5, Benjamin B Machta6,2
1Computational Biology and Bioinformatics Program, Yale University, New Haven, CT-06520, USA.
Biorxiv : the Preprint Server for Biology
|November 28, 2024
Summary
Researchers developed an AI algorithm to map evolutionary paths between proteins. This tool helps understand protein evolution and homology by generating viable intermediate sequences using the ESM2 protein language model.
Area of Science:
- Computational Biology
- Protein Engineering
- Evolutionary Bioinformatics
Background:
- Protein structure and function are dictated by amino acid sequences, shaped by evolutionary forces.
- Neutral networks connect local protein sequence spaces, but large-scale morphospace connectivity is unclear.
Purpose of the Study:
- To develop an algorithm for generating viable evolutionary paths between distantly related proteins.
- To explore the large-scale connectedness of protein morphospace using AI-driven methods.
Main Methods:
- Utilized artificial intelligence tools for protein structure prediction and functional plausibility assessment.
- Developed a novel algorithm employing the ESM2 protein language model to evaluate sequence fitness.
- Generated intermediate protein sequences via single residue substitutions, insertions, and deletions.
Main Results:
- Successfully generated viable paths between diverse protein pairs, including those with different structural folds.
- Documented qualitative variations in paths connecting progressively divergent protein sequences.
- Demonstrated the potential of sequence interpolation as a proxy for homology detection.
Conclusions:
- The developed algorithm provides insights into protein evolutionary trajectories and morphospace.
- Sequence interpolation ease can serve as a quantitative measure for assessing protein homology.
- This approach advances our understanding of protein evolution and sequence-structure-function relationships.
Related Concept Videos
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Protein Complexes with Interchangeable Parts
2.5K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.5K
Protein-Protein Interfaces
3.7K
3.7K


