Related Experiment Videos
Experiments in searching small proteins in unannotated large eukaryotic genomes
Jacques Colinge1, Isabelle Cusin, Samia Reffas
1GeneProt Inc., 2 Pré-de-la-Fontaine, CH-1217 Meyrin, Switzerland. jcolinge@excite.com
Journal of Proteome Research
|February 15, 2005
Summary
Directly searching genome sequences with mass spectrometry data aids genome annotation. This study addresses challenges in identifying peptides spanning exon/intron boundaries in large genomes, improving gene structure prediction.
Area of Science:
- Proteomics
- Genomics
- Bioinformatics
Background:
- Mass spectrometry data is increasingly used for direct genome sequence searching.
- This approach can correct and complement existing genome annotations.
- Searching large eukaryotic genomes with peptide tandem mass spectra presents practical difficulties, especially for small proteins (<40 kDa).
Purpose of the Study:
- To explore the challenges of automatically identifying peptides that span across exon/intron boundaries using experimental data.
- To assess the impact of mass spectrometry data on genome annotation accuracy.
- To demonstrate the utility of proteomics data in refining predicted gene structures.
Main Methods:
- Directly searching large eukaryotic genome sequences using peptide ion trap tandem mass spectra.
- Analyzing peptides that span across exon/intron boundaries.
- Evaluating search results against Swiss-Prot for comparison.
- Assessing the effect of parent mass accuracy on peptide identification rates.
Main Results:
- Approximately 30% of peptides were missed in a human genome search compared to a Swiss-Prot search.
- This peptide miss rate was significantly reduced with improved parent mass accuracy.
- Several examples of predicted gene structures were identified that could benefit from proteomics data, particularly peptides spanning exon/intron boundaries.
Conclusions:
- Directly searching genome sequences with mass spectrometry data is a valuable approach for improving genome annotations.
- Identifying peptides spanning exon/intron boundaries is a key challenge but crucial for accurate gene structure prediction.
- Enhanced mass accuracy in proteomics data significantly improves the identification of relevant peptides, aiding in the refinement of gene models.