Related Experiment Videos
SPIDER: software for protein identification from sequence tags with de novo sequencing error
Yonghua Han1, Bin Ma, Kaizhong Zhang
1Department of Computer Science, University of Western Ontario, London, Canada. yhan2@csd.uwo.ca
Summary
We developed SPIDER, an efficient algorithm and software for identifying proteins and peptides by matching error-containing sequence tags from MS/MS data to protein databases.
Area of Science:
- Proteomics
- Bioinformatics
- Computational Biology
Background:
- De novo sequencing of MS/MS spectra generates sequence tags for protein identification.
- Accurate sequence tags enable protein homolog identification using tools like MS-BLAST.
- De novo sequencing frequently yields partially incorrect sequence tags, hindering protein identification.
Purpose of the Study:
- To develop an efficient algorithm for matching error-containing sequence tags to database sequences.
- To create a software package, SPIDER, for improved protein and peptide identification.
- To provide a free, publicly accessible tool for researchers in proteomics and bioinformatics.
Main Methods:
- Developed a novel algorithm to efficiently match sequence tags with common errors (e.g., segment replacement) to protein databases.
- Implemented the algorithm into a user-friendly software package named SPIDER.
- Made the SPIDER software freely available via the internet for public use.
Main Results:
- The SPIDER algorithm effectively matches sequence tags containing errors to database sequences.
- SPIDER facilitates more accurate and comprehensive protein and peptide identification from MS/MS data.
- The software package offers an efficient solution for a common challenge in de novo sequencing.
Conclusions:
- SPIDER provides an efficient and effective method for protein and peptide identification using error-prone sequence tags.
- The developed algorithm and software address a significant limitation in de novo sequencing for proteomics.
- SPIDER is a valuable, freely accessible resource for the scientific community involved in mass spectrometry-based protein identification.