Protein embedding based alignment
Benjamin Giovanni Iovino1, Yuzhen Ye2
1Luddy School of Informatics, Computing and Engineering, Indiana University, 700 N. Woodlawn Avenue, Bloomington, IN, 47408, USA.
BMC Bioinformatics
|February 27, 2024
Summary
Protein Embedding based Alignments (PEbA) improves protein sequence alignment in the difficult twilight zone. This new method, using protein language model embeddings, outperforms traditional substitution matrices and other recent alignment tools.
Area of Science:
- Bioinformatics
- Computational Biology
- Protein Sequence Analysis
Background:
- Traditional protein sequence alignment methods struggle with sequences below 35% identity.
- Substitution matrices, developed in the 1970s, are suboptimal for scoring alignments in the low-identity 'twilight zone'.
Purpose of the Study:
- To develop a novel algorithm, Protein Embedding based Alignments (PEbA), for improved protein sequence alignment.
- To leverage protein language models for enhanced alignment accuracy in the low-identity twilight zone.
Main Methods:
- PEbA utilizes a dynamic programming approach similar to Smith-Waterman.
- Amino acid matching scores in PEbA are derived from the similarity of embeddings generated by protein language models.
- The algorithm was evaluated on over 12,000 benchmark pairwise alignments from BAliBASE.
Main Results:
- PEbA significantly outperformed traditional BLOSUM substitution matrix-based alignments, especially for sequences with <10% identity (over fourfold improvement).
- Comparisons with different protein language models showed ProtT5-XL-U50 embeddings yielded the best alignment performance.
- PEbA demonstrated superior performance compared to other recent embedding-based alignment methods like DEDAL and vcMSA.
Conclusions:
- General-purpose protein language models provide valuable contextual information for protein sequence alignment.
- PEbA offers a more accurate alignment method than traditional approaches, particularly for divergent protein sequences.
Related Concept Videos
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein-Protein Interfaces
3.8K
3.8K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Lipids as Anchors
5.6K
In the plasma membrane, the lipids forming the bilayer can also act as an anchor to tether proteins to the membrane. The three main types of lipid anchors found in eukaryotes are – prenyl groups, fatty acyl groups, and glycosylphosphatidylinositol or GPI groups. Prenyl and fatty acyl groups act as anchors on the cytosolic surface of the membrane, whereas GPI anchors proteins on the extracellular side.
The carboxy-terminal of most of the prenylated proteins, such as Ras proteins, contains...
The carboxy-terminal of most of the prenylated proteins, such as Ras proteins, contains...
5.6K
Tail-anchoring of Proteins in the ER Membrane
3.1K
Tail-anchored, or TA, proteins are estimated to make up to 3-5% of membrane proteins found in the eukaryotic cell. Such proteins have a single transmembrane domain located approximately 30 amino acid residues upstream from the C-terminal end. As a result, the signal recognition particle (SRP) cannot guide a TA protein to the ER membrane for cotranslational insertion. Hence, they are integrated into the ER membrane post-translationally using their C-terminal end as the anchor. TA proteins...
3.1K
Protein Complex Assembly
10.6K
Proteins can form homomeric complexes with another unit of the same protein or heteromeric complexes with different types. Most protein complexes self-assemble spontaneously via ordered pathways, while some proteins need assembly factors that guide their proper assembly. Despite the crowded intracellular environment, proteins usually interact with their correct partners and form functional complexes.
Many viruses self-assemble into a fully functional unit using the infected host cell to...
Many viruses self-assemble into a fully functional unit using the infected host cell to...
10.6K


