Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genome Annotation and Assembly03:36

Genome Annotation and Assembly

The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
RNA-seq03:21

RNA-seq

RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Conservation of Protein Domains Over Different Proteins02:26

Conservation of Protein Domains Over Different Proteins

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Tagging and Fusion Proteins01:24

Tagging and Fusion Proteins

Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
Conservation of Protein Domains02:26

Conservation of Protein Domains

Protein domains are small structurally independent units that are part of a single amino acid chain.  Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites01:49

Conserved Binding Sites

Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Survival prediction from neural parametrization of diffusive processes.

Physical review. E·2026
Same author

Update of the MSKCC nomogram for metastatic progression and its role in active surveillance: the Italian TPCP cohort.

Frontiers in oncology·2026
Same author

Environmental Personal Exposure Clusters to Investigate Multiple Sclerosis and Amyotrophic Lateral Sclerosis Progression.

Studies in health technology and informatics·2026
Same author

On the state of protein function prediction: a report on the fourth CAFA challenge.

bioRxiv : the preprint server for biology·2026
Same author

Advances in Protein Function Prediction from the Fifth CAFA Challenge.

bioRxiv : the preprint server for biology·2026
Same author

A machine learning-derived cardiovascular risk score in people with HIV: the ML-ICONA score.

American journal of preventive cardiology·2026

Related Experiment Video

Updated: May 13, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
08:03

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations

Published on: December 7, 2021

How to inherit statistically validated annotation within BAR+ protein clusters.

Damiano Piovesan1, Pier Luigi Martelli, Piero Fariselli

  • 1Bologna Biocomputing Group, University of Bologna, Italy.

BMC Bioinformatics
|March 22, 2013
PubMed
Summary

This study validates transferring protein annotation, including Gene Ontology (GO) terms and Pfam domains, to new sequences. It proves that even distantly related proteins can be accurately annotated with statistical validation.

More Related Videos

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
10:12

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues

Published on: January 10, 2019

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

Related Experiment Videos

Last Updated: May 13, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
08:03

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations

Published on: December 7, 2021

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
10:12

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues

Published on: January 10, 2019

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Genomics

Background:

  • Protein annotation is crucial for understanding biological function.
  • Current methods rely heavily on sequence similarity, raising questions about accuracy for distantly related proteins.
  • UniProtKB is the primary resource for protein sequence and functional information.

Purpose of the Study:

  • To develop a statistically validated method for protein sequence annotation.
  • To assess the transferability of structural and functional features from known proteins to new sequences.
  • To address the challenge of annotating proteins with low sequence homology to known templates.

Main Methods:

  • Utilized the Critical Assessment of Function Annotations (CAFA) dataset of 48,298 proteins.
  • Developed BAR+, an annotation resource for discriminating statistically validated annotations.
  • Aligned CAFA sequences against BAR+ for annotation transfer and validation.

Main Results:

  • Successfully transferred Gene Ontology (GO) terms to 68% and Pfam domains to 72% of CAFA sequences.
  • Validated existing annotations for 78% of the CAFA set.
  • Assigned new, statistically validated annotations to 14.8% of sequences and identified new structural templates for 25% of chains, even with <30% sequence identity.

Conclusions:

  • Statistical validation enables safe annotation transfer, even for distantly related homologs.
  • Clustering protein sequences and validating shared features within clusters is key to reliable annotation.
  • This approach enhances the accuracy and scope of protein functional and structural annotation.