Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Protein Networks02:26

Protein Networks

3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Proteomics01:33

Proteomics

7.3K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.3K
Protein Families02:47

Protein Families

15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism.   Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members.   If these new proteins contain similar amino acids in key...
15.3K
Protein Organization01:24

Protein Organization

6.4K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
6.4K
Protein-protein Interfaces02:04

Protein-protein Interfaces

12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A novel method to select Reference Proteomes in UniProt.

bioRxiv : the preprint server for biology·2026
Same author

Advances in Protein Function Prediction from the Fifth CAFA Challenge.

bioRxiv : the preprint server for biology·2026
Same author

Expanding the human proteome with microproteins and peptideins.

Nature·2026
Same author

From use cases to infrastructure: a cross-institutional survey of priorities in data-driven biomedical research.

Journal of the American Medical Informatics Association : JAMIA·2026
Same author

The PanOryza pangene catalog of Asian cultivated rice.

Genome research·2025
Same author

EMBL's European Bioinformatics Institute (EMBL-EBI) in 2025.

Nucleic acids research·2025

Related Experiment Video

Updated: Jun 24, 2025

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
07:38

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames

Published on: April 11, 2019

12.7K

Improved selection of canonical proteins for reference proteomes.

Giuseppe Insana1, Maria J Martin1, William R Pearson2

  • 1European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton CB10 1SD, UK.

NAR Genomics and Bioinformatics
|June 12, 2024
PubMed
Summary

UniProt canonical protein sequences can be inconsistent in higher eukaryotes. The ortho2tree pipeline improves canonical assignment by identifying biologically relevant isoforms, enhancing protein annotation accuracy.

More Related Videos

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

7.1K
An Integrated Approach for Microprotein Identification and Sequence Analysis
09:37

An Integrated Approach for Microprotein Identification and Sequence Analysis

Published on: July 12, 2022

3.4K

Related Experiment Videos

Last Updated: Jun 24, 2025

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
07:38

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames

Published on: April 11, 2019

12.7K
Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

7.1K
An Integrated Approach for Microprotein Identification and Sequence Analysis
09:37

An Integrated Approach for Microprotein Identification and Sequence Analysis

Published on: July 12, 2022

3.4K

Area of Science:

  • Bioinformatics
  • Proteomics
  • Genomics

Background:

  • UniProt canonical protein sequences are crucial for research but can be inconsistent in higher eukaryotes due to isoform selection.
  • The longest sequence is often chosen as canonical in unreviewed protein databases, leading to biologically unlikely length variations in highly similar orthologs.

Purpose of the Study:

  • To develop and validate the ortho2tree pipeline for improving canonical protein sequence assignment.
  • To address inconsistencies in canonical protein selection arising from alternative splicing in higher eukaryotes.

Main Methods:

  • The ortho2tree pipeline analyzes orthologous protein sequences, builds multiple alignments, and constructs phylogenetic trees to identify isoforms with similar lengths.
  • It examines canonical and isoform sequences from UniProt Reference Proteomes, focusing on mammals.

Main Results:

  • ortho2tree proposed 7804 canonical changes and confirmed 53,434 canonicals in UniProtKB release 2023_01.
  • The pipeline's isoform selection resulted in gap distributions comparable to those in bacteria and yeast, suggesting improved biological accuracy.
  • ortho2tree showed high agreement with the MANE (Matched Annotation للنظام Eukaryotes) standard, with 82% agreement for proposed changes and 92% for confirmed canonicals.

Conclusions:

  • The ortho2tree pipeline offers a more accurate method for assigning canonical protein sequences, particularly in vertebrates and plants.
  • Improved canonical assignment enhances the reliability of protein similarity searching, functional annotation, and structural analysis.