Refining Transcriptome Gene Catalogs by MS-Validation of Expressed Proteins
Sirius P K Tse1,2, Mathieu Beauchemin3, David Morse3
1Shenzhen Key Laboratory of Food Biological Safety Control, The Hong Kong Polytechnic University, Kowloon, Hong Kong.
Identifying protein sequences using mass spectrometry (LC-MS/MS) in dinoflagellates reveals the actual protein-coding frames. This MS-validated protein data offers a more accurate view of cellular function than transcriptomes alone.
Area of Science:
- Proteomics
- Molecular Biology
- Bioinformatics
Background:
- Tandem mass spectrometry (LC-MS/MS) identifies proteins in complex mixtures, aiding biological function studies.
- Transcriptomes are crucial for non-model organisms, serving as gene catalogs and revealing metabolic potential.
- Determining the correct protein-coding reading frame is essential for accurate biological interpretation.
Purpose of the Study:
- To identify the actively translated protein-coding reading frames within the transcriptome of the dinoflagellate Lingulodinium polyedra.
- To assess the impact of using MS-validated protein sequences versus nucleic acid sequences on downstream bioinformatics analyses.
- To compare MS-validated protein sequences with longest open reading frame (ORF) datasets and evaluate protein-mRNA abundance correlations.
Main Methods:
- Combined LC-MS/MS data from multiple experiments on Lingulodinium polyedra extracts.
- Utilized a 74,655-sequence transcriptome to identify peptide matches.
- Compiled a dataset of 6,628 MS-validated protein sequences and compared them with their parental nucleic acid sequences using BLASTp and BLASTx.
Main Results:
- MS-validated protein sequences showed differences in gene ontology, BLAST hits, and KEGG pathway enzyme content compared to DNA sequence BLASTx analyses.
- MS-validated protein sequences differed from longest open reading frame (ORF) datasets.
- A poor correlation was observed between protein and mRNA abundance levels, a novel finding for dinoflagellates.
Conclusions:
- MS-validated protein sequence datasets provide a more accurate representation of cellular capacity than nucleic acid sequence datasets.
- Using MS-validated protein sequences can enhance the accuracy of downstream analyses in proteomics and transcriptomics.
- Developing MS-validated protein datasets can accelerate the interpretation of MS-MS spectra in proteomics.
More Related Videos
10:37Deep Proteome Profiling by Isobaric Labeling, Extensive Liquid Chromatography, Mass Spectrometry, and Software-assisted Quantification
Published on: November 15, 2017
14:51Comprehensive Workflow of Mass Spectrometry-based Shotgun Proteomics of Tissue Samples
Published on: November 13, 2021
