Related Experiment Videos
SPIDER: software for protein identification from sequence tags with de novo sequencing error
Yonghua Han1, Bin Ma, Kaizhong Zhang
1Department of Computer Science, University of Western Ontario, London, Ontario, Canada N6A 5B7, Canada. yhan2@csd.uwo.ca
Journal of Bioinformatics and Computational Biology
|August 19, 2005
Summary
We developed SPIDER, an efficient algorithm and software for identifying proteins using mass spectrometry/mass spectrometry (MS/MS) sequence tags, even with common sequencing errors. This tool aids in accurate protein and peptide identification.
Area of Science:
- Proteomics
- Bioinformatics
- Computational Biology
Background:
- Mass spectrometry/mass spectrometry (MS/MS) is crucial for identifying novel proteins.
- De novo sequencing generates sequence tags from MS/MS spectra for database matching.
- Existing methods struggle with partially correct sequence tags, often due to mass-equivalent amino acid substitutions.
Purpose of the Study:
- To develop an efficient algorithm for matching error-containing sequence tags to database sequences.
- To introduce SPIDER, a software package for improved protein and peptide identification.
- To provide a publicly accessible tool for the scientific community.
Main Methods:
- Developed a novel algorithm to handle sequence tags with common errors, such as segment-for-segment substitutions.
- Implemented the algorithm into a user-friendly software package named SPIDER.
- Made the SPIDER software freely available online.
Main Results:
- The SPIDER algorithm efficiently matches sequence tags with errors to protein databases.
- SPIDER improves the accuracy of protein and peptide identification compared to methods relying on perfectly correct tags.
- The software package facilitates the identification of homologs even when de novo sequencing yields imperfect data.
Conclusions:
- SPIDER offers an effective solution for protein and peptide identification challenges posed by de novo sequencing errors.
- The developed algorithm and software enhance the reliability of MS/MS-based protein identification.
- SPIDER is a valuable, free resource for researchers in proteomics and bioinformatics.