Related Experiment Videos
Randomized sequence databases for tandem mass spectrometry peptide and protein identification.
Roger Higdon1, Jason M Hogan, Gerald Van Belle
1The BIATECH Institute, 19310 N. Creek Parkway South, Suite 115, Bothell, WA 98011, USA.
Omics : a Journal of Integrative Biology
|January 13, 2006
Summary
Accurately identifying peptides and proteins using tandem mass spectrometry (MS/MS) requires study-specific methods. This study recommends combined searches of reshuffled and forward databases to quantify false positive rates, ensuring reliable protein identification.
Area of Science:
- Proteomics
- Bioinformatics
- Analytical Chemistry
Background:
- Tandem mass spectrometry (MS/MS) with database searching is standard for high-throughput peptide and protein identification.
- Current accuracy assessments like "high confidence" lack specificity and vary with experimental conditions.
- There is a need for robust, study-specific methods to estimate false positive rates in MS/MS identifications.
Purpose of the Study:
- To evaluate methods for estimating false positive identification rates in MS/MS data using randomized databases.
- To compare error rate estimations from randomized database searches with actual error rates from known protein standards.
- To provide a quantitative measure of peptide and protein identification accuracy.
Main Methods:
- Searches were performed using forward and randomized (reversed and reshuffled) databases.
- Separate and combined search strategies were examined.
- Methods were validated using MS/MS runs of known protein standards and applied to Shewanella oneidensis MR-1 samples.
Main Results:
- Randomized database searches provide quantitative estimates of false positive rates.
- Combined searches of a reshuffled database appended to a forward database were effective in estimating error rates.
- The proposed method allows for setting specific error rate thresholds.
Conclusions:
- Combined searches of reshuffled and forward databases are recommended for quantifying false positive rates in peptide and protein identification.
- This approach offers direct and quantifiable measures of accuracy, replacing vague assessments.
- Enables researchers to achieve desired error rates and enhances the reliability of proteomic data.