Related Experiment Video
Updated: Jun 12, 2026

09:16
A Spin-Tip Enrichment Strategy for Simultaneous Analysis of N-Glycopeptides and Phosphopeptides from Human Pancreatic Tissues
Published on: May 4, 2022
An efficient parallelization of phosphorylated peptide and protein identification
Leheng Wang1, Wenping Wang, Hao Chi
1Key Lab of Intelligent Information Processing, Chinese Academy of Sciences, Beijing 100190, P.R. China.
Rapid Communications in Mass Spectrometry : RCM
|May 26, 2010
Summary
This study introduces a computational model to optimize parallel computing for protein identification using tandem mass spectrometry. The developed scheduling methods significantly accelerate proteomics data analysis, reducing processing time from hours to minutes.
Area of Science:
- Computational Biology
- Proteomics
- Bioinformatics
Background:
- Protein identification via tandem mass spectrometry is crucial for biological research.
- Increasing computational demands necessitate efficient data analysis techniques.
- Parallel computing offers a viable solution for accelerating proteomics workflows.
Purpose of the Study:
- To investigate factors influencing the runtime of the pFind search engine.
- To develop an estimation model for predicting and optimizing parallel search engine performance.
- To create effective on-line and off-line scheduling strategies for parallel proteomics data analysis.
Main Methods:
- Developed a runtime estimation model for the pFind search engine.
- Implemented novel on-line and off-line scheduling algorithms based on the model.
- Evaluated performance using public phosphopeptide datasets (PhosphoPep) and a large-scale spectra dataset.
Main Results:
- Achieved significant speedups: 83.7x on 100 processors for phosphopeptides.
- Reduced identification time from over 10 hours to 9 minutes on a single PC.
- Demonstrated high efficiency (80.9%) with a 258.9x speedup on 320 processors for a large dataset.
Conclusions:
- The developed estimation model and scheduling methods effectively accelerate protein identification.
- Parallel computing significantly enhances the efficiency of proteomics data analysis.
- The approach is applicable to other protein sequence search engines.

