Improved selection of canonical proteins for reference proteomes
Giuseppe Insana1, Maria J Martin1, William R Pearson2
1European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton CB10 1SD, UK.
NAR Genomics and Bioinformatics
|June 12, 2024
Summary
UniProt canonical protein sequences can be inconsistent in higher eukaryotes. The ortho2tree pipeline improves canonical assignment by identifying biologically relevant isoforms, enhancing protein annotation accuracy.
Area of Science:
- Bioinformatics
- Proteomics
- Genomics
Background:
- UniProt canonical protein sequences are crucial for research but can be inconsistent in higher eukaryotes due to isoform selection.
- The longest sequence is often chosen as canonical in unreviewed protein databases, leading to biologically unlikely length variations in highly similar orthologs.
Purpose of the Study:
- To develop and validate the ortho2tree pipeline for improving canonical protein sequence assignment.
- To address inconsistencies in canonical protein selection arising from alternative splicing in higher eukaryotes.
Main Methods:
- The ortho2tree pipeline analyzes orthologous protein sequences, builds multiple alignments, and constructs phylogenetic trees to identify isoforms with similar lengths.
- It examines canonical and isoform sequences from UniProt Reference Proteomes, focusing on mammals.
Main Results:
- ortho2tree proposed 7804 canonical changes and confirmed 53,434 canonicals in UniProtKB release 2023_01.
- The pipeline's isoform selection resulted in gap distributions comparable to those in bacteria and yeast, suggesting improved biological accuracy.
- ortho2tree showed high agreement with the MANE (Matched Annotation للنظام Eukaryotes) standard, with 82% agreement for proposed changes and 92% for confirmed canonicals.
Conclusions:
- The ortho2tree pipeline offers a more accurate method for assigning canonical protein sequences, particularly in vertebrates and plants.
- Improved canonical assignment enhances the reliability of protein similarity searching, functional annotation, and structural analysis.
Related Concept Videos
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Proteomics
7.3K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.3K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Protein Organization
6.4K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.4K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K


