Improving prediction of protein subcellular localization using evolutionary information and sequence-order

Minghui Wang1, Ao Li, Dan Xie

  • 1Department of Electronic Science and Technology, University of Science and Technology of China, Hefei, Anhui 230026, China.

Summary

Accurate prediction of protein subcellular localization is crucial for understanding biological function. This study introduces a novel hybrid method using evolutionary and sequence data, achieving high performance in predicting protein locations.

Related Concept Videos

Signal Sequences and Sorting Receptors01:41

Signal Sequences and Sorting Receptors

Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
Evolutionary Relationships through Genome Comparisons02:54

Evolutionary Relationships through Genome Comparisons

Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Conserved Binding Sites01:49

Conserved Binding Sites

Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Families02:47

Protein Families

Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key locations, protein...
Directing Proteins to the Rough Endoplasmic Reticulum01:34

Directing Proteins to the Rough Endoplasmic Reticulum

The organelle-specific signaling sequences direct proteins synthesized in the cytosol to their final destination like ER, mitochondria, peroxisomes, etc. Some of the proteins directed to ER are then trafficked via vesicles to other organelles within the cell or the extracellular environment through the Golgi complex. For example, the rough ER synthesizes soluble proteins for transportation to the lysosomes or secretion out of the cell. It can also synthesize transmembrane proteins that can...