Spectral Prediction Features as a Solution for the Search Space Size Problem in Proteogenomics
Steven Verbruggen1, Siegfried Gessulat2, Ralf Gabriels3
1BioBix, Lab of Bioinformatics and Computational Genomics, Department of Mathematical Modeling, Statistics and Bioinformatics, Faculty of Bioscience Engineering, Ghent University, Ghent, Belgium; OHMX.bio, Ghent, Belgium.
Molecular & Cellular Proteomics : MCP
|April 6, 2021
Summary
Combining spectral intensity predictors with MaxQuant scores improves peptide identification and validation in proteogenomics. This approach enhances confidence in matching peptides to spectra, even with large databases.
Area of Science:
- Proteomics
- Bioinformatics
- Computational Biology
Background:
- Proteogenomics aims to link genomic information with protein expression data.
- Distinguishing true peptide-spectrum matches from false positives is challenging, especially with large protein databases.
- Tandem mass spectrometry (MS/MS) is crucial for peptide identification, but accuracy can be limited.
Purpose of the Study:
- To enhance peptide identification rates and validation stringency in proteogenomics.
- To evaluate the utility of spectral intensity predictors in improving peptide-to-spectrum matching.
- To integrate features from MS2PIP and Prosit with existing scoring methods for proteogenomic analysis.
Main Methods:
- Utilized protein sequence databases generated from ribosome profiling and nanopore RNA-Seq data.
- Extracted features from MS2PIP and Prosit, tandem mass spectrometry intensity predictors.
- Combined these features with canonical scores from MaxQuant within the Percolator postprocessing tool.
- Applied the integrated approach to assess peptide identification and validation.
Main Results:
- Demonstrated that incorporating spectral intensity predictor features significantly enhances peptide identification rates.
- Showed improved validation stringency for peptide-to-spectrum matches.
- Confirmed the effectiveness of the combined approach for proteogenomic datasets.
Conclusions:
- The integration of MS2PIP and Prosit features with MaxQuant scores in Percolator offers a robust method for improving proteogenomic analyses.
- This strategy effectively addresses the challenge of distinguishing true from false peptide-spectrum matches.
- The findings support the use of spectral intensity prediction features for more confident and comprehensive proteogenomic discoveries.
Related Concept Videos
Proteomics
8.7K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
8.7K
Peptide Identification Using Tandem Mass Spectrometry
7.5K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
7.5K


