A two-step database search method improves sensitivity in peptide sequence matches for metaproteomics and
Pratik Jagtap1, Jill Goslinga, Joel A Kooren
1Minnesota Supercomputing Institute, Minneapolis, MN, USA. pratik@msi.umn.edu
Proteomics
|February 16, 2013
Summary
A new two-step method significantly improves peptide matching in large-scale metaproteomics and proteogenomics. This approach doubles high-confidence peptide matches by reducing false negatives, enhancing protein identification sensitivity.
Area of Science:
- Biochemistry
- Bioinformatics
- Proteomics
Background:
- Metaproteomic and proteogenomic studies utilize large sequence databases (>10^6 sequences).
- Database searching of MS/MS data against these large databases faces challenges in peptide sequence matching.
- Strict filtering to minimize false positives often results in an increase of false negatives, limiting peptide identification.
Purpose of the Study:
- To develop and validate a novel two-step method for enhanced peptide sequence matching in large databases.
- To improve the sensitivity and accuracy of protein identification in metaproteomics and proteogenomics.
- To overcome the limitations of conventional one-step database search methods.
Main Methods:
- A primary database search against a large sequence database was conducted.
- A smaller subset database was constructed using the initial matches.
- A second search was performed against a target-decoy version of the subset database merged with a host database.
- High-confidence peptide matches were used for protein inference.
Main Results:
- The two-step method yielded approximately double the number of high-confidence peptide matches in both metaproteomic and proteogenomic analyses compared to the one-step method.
- The majority of additional matches identified by the two-step method were previously classified as false negatives by the one-step method.
- The improved peptide matching sensitivity was consistent across different database search programs.
- The method effectively captured nearly all peptides identified by the conventional approach.
Conclusions:
- The developed two-step method significantly enhances peptide matching sensitivity for large-scale metaproteomics and proteogenomics.
- This approach is particularly valuable for maximizing protein identification in complex biological samples.
- The method offers a robust solution to the challenge of false negatives in large database searches.
Related Concept Videos
Peptide Identification Using Tandem Mass Spectrometry
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Proteomics
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term proteomics...


