Representative proteomes: a stable, scalable and unbiased proteome set for sequence analysis and functional
Chuming Chen1, Darren A Natale, Robert D Finn
1Center for Bioinformatics and Computational Biology, University of Delaware, Newark, Delaware, United States of America.
Plos One
|May 11, 2011
Summary
We developed Representative Proteomes (RPs) to reduce protein sequence data by over 80% while preserving essential information. This approach aids computational analysis and protein characterization, offering diverse granularity options.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Exponential growth in protein sequences challenges computational and manual analysis resources.
- Minimizing data while retaining information is crucial for efficient biological data handling.
Purpose of the Study:
- To develop a method for reducing the size of proteomic datasets.
- To create a curated set of Representative Proteomes (RPs) that capture maximal information from similar proteomes.
- To provide RPs at various granularity levels for flexible data analysis.
Main Methods:
- Proteomes were clustered into Representative Proteome Groups (RPGs) based on UniRef50 co-membership.
- Representative Proteomes (RPs) were selected to best represent their respective RPGs.
- RPs were generated using different co-membership thresholds (CMT): 75%, 55%, 35%, and 15%.
Main Results:
- A 55% co-membership threshold (RP55) demonstrated strong alignment with standard taxonomic classifications.
- The RP set reduced sequence space by over 80% compared to UniProtKB.
- Sequence diversity (over 95% of InterPro domains) and annotation information (93% of experimentally characterized proteins) were retained.
Conclusions:
- Representative Proteomes offer a significant reduction in data size while maintaining high information content.
- RPs facilitate efficient sequence similarity searches, protein classification, and targeted characterization.
- The RP dataset provides a valuable resource for managing and analyzing large-scale proteomic data.
Related Concept Videos
Proteomics
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Peptide Identification Using Tandem Mass Spectrometry
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Ribosome Profiling
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique helps...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...


