Related Experiment Video
Updated: Aug 11, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Defining parameters for homology-tolerant database searching
J P Kayser1, J L Vallet, R L Cerny
1USDA, ARS, RLH US Meat Animal Research Center, Clay Center, NE 68933, USA.
De novo sequencing with tandem mass spectrometry (MS/MS) enables protein identification. A homology search strategy using 20 peptides and 30% mismatch optimizes protein database searching when sequence data is limited.
Area of Science:
- Proteomics
- Bioinformatics
- Mass Spectrometry
Background:
- De novo interpretation of tandem mass spectrometry (MS/MS) spectra is crucial for protein database searching with limited sequence information.
- Developing effective strategies for homology-tolerant database searches is essential for accurate protein identification.
Purpose of the Study:
- To define an optimal strategy for homology-tolerant database searching using de novo MS/MS spectra.
- To compare the efficacy of homology searches with traditional peptide mass searching (Mascot).
Main Methods:
- Homology searches were performed using MS-Homology software with varying numbers of abundant peptides (20, 10, or 5) and mismatch tolerances (50%, 30%, 10%).
- Protein scores were corrected against a threshold derived from random peptides.
- Protein identification results from homology searches were compared to those obtained using Mascot with MS/MS data.
Main Results:
- The optimal homology search strategy involved submitting 20 peptides with a 30% mismatch tolerance, yielding the highest corrected protein scores (p < .01).
- Mascot and homology searches (using 20 or all peptides) identified the same top-ranking protein for 63.4% of analyzed spots.
- Mascot achieved greater percent coverage (25.1%) compared to homology searches using all (18.3%) or the top 20 peptides (10.6%).
- De novo sequences showed complete matches to known sequences in 35% of cases, increasing to 44.0% when limited to the top 20 peptides.
Conclusions:
- A homology search strategy using 20 peptides and 30% mismatch is effective for protein identification from de novo MS/MS sequences.
- While Mascot offers higher percent coverage, homology searches are valuable for database searching with limited sequence data.
- Combining MS-Homology for initial identification with a subsequent peptide mass search can enhance overall protein coverage.
More Related Videos
07:49Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
Related Concept Videos
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Protein Families
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Evolutionary Relationships through Genome Comparisons
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...