High fitness paths can connect proteins with low sequence overlap
Pranav Kantroo1,2, Günter P Wagner3,4,5, Benjamin B Machta6,2
1Computational Biology and Bioinformatics Program, Yale University, New Haven, CT-06520, USA.
Arxiv
|November 28, 2024
Summary
Researchers developed an AI algorithm to map evolutionary paths between proteins. This tool explores protein sequence space, aiding in understanding protein relationships and homology.
Area of Science:
- Computational Biology
- Protein Engineering
- Bioinformatics
Background:
- Protein structure and function are dictated by amino acid sequences.
- Evolutionary forces shape protein folds and activities, with neutral networks connecting local sequence spaces.
- The broader connectivity of protein morphospace is not well understood.
Purpose of the Study:
- To develop an algorithm for generating viable evolutionary paths between distantly related proteins.
- To leverage artificial intelligence for predicting protein structure and functional plausibility.
- To explore the large-scale connectedness of protein morphospace.
Main Methods:
- An algorithm was developed using AI tools to predict protein structure and functional plausibility.
- Viable paths between protein pairs were generated using single residue substitutions, insertions, and deletions.
- The protein language model ESM2 was used to evaluate sequence fitness during path generation.
Main Results:
- The algorithm successfully generated viable paths between progressively divergent protein pairs.
- Qualitative variations were observed across generated paths, including those between proteins with different structural folds.
- The ease of interpolation between sequences can serve as a proxy for homology.
Conclusions:
- The developed algorithm effectively navigates protein sequence space, connecting distantly related proteins.
- This approach provides insights into protein evolution and the underlying principles of protein morphospace.
- The method offers a novel way to assess the likelihood of homology between protein sequences.
Related Concept Videos
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Protein Complexes with Interchangeable Parts
2.5K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.5K
Protein-Protein Interfaces
3.7K
3.7K


