Related Experiment Videos
A statistical model for identifying proteins by tandem mass spectrometry
Alexey I Nesvizhskii1, Andrew Keller, Eugene Kolker
1Institute for Systems Biology, 1441 North 34th Street, Seattle, Washington 98103, USA. nesvi@systemsbiology.org
Analytical Chemistry
|November 25, 2003
Summary
This study introduces a statistical model to accurately determine protein presence probabilities from mass spectrometry data. The method enhances the reliability of large-scale proteomics by providing predictable error rates and a standardized approach for data analysis.
Area of Science:
- Proteomics and Bioinformatics
- Statistical Modeling in Mass Spectrometry
Background:
- Accurate protein identification from mass spectrometry (MS/MS) data is crucial for biological research.
- Existing methods face challenges in handling peptides assigned to multiple proteins and quantifying identification reliability.
Purpose of the Study:
- To develop a statistical model for calculating protein presence probabilities in biological samples.
- To establish a reliable method for filtering large-scale proteomics datasets with controlled false positive rates.
Main Methods:
- A statistical model was developed to compute protein presence probabilities based on peptide assignments to MS/MS spectra.
- Peptides mapping to multiple proteins were apportioned, and the expectation-maximization algorithm identified a minimal protein list.
- The model was validated using purified proteins and complex biological samples (H. influenzae, Halobacterium).
Main Results:
- The model accurately computes protein presence probabilities, demonstrating high power in distinguishing correct from incorrect identifications.
- Demonstrated ability to filter large-scale proteomics data with predictable sensitivity and false positive error rates.
- The approach is fast, consistent, and transparent, facilitating reproducible large-scale proteomics studies.
Conclusions:
- The presented statistical model offers an accurate and powerful method for protein identification in proteomics.
- This approach provides a standardized framework for publishing and comparing large-scale proteomics data.
- Enables reliable filtering of proteomics datasets, improving the quality and interpretability of results.