Related Experiment Video
Updated: Jul 19, 2026

Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
Published on: April 11, 2019
MASCOT HTML and XML parser: an implementation of a novel object model for protein identification data
Chunguang G Yang1, Stephen J Granite, Jennifer E Van Eyk
1Center for Cardiovascular Bioinformatics and Modeling, The Institute for Computational Medicine and The Whitaker Biomedical Engineering Institute, The Johns Hopkins University, Baltimore, MD 21218-2686, USA.
A new Protein Identification Data Object Model (PDOM) and parser streamline proteomics data analysis. This tool simplifies storing and analyzing MASCOT search results from mass spectrometry (MS).
Area of Science:
- Proteomics
- Bioinformatics
- Computational Biology
Background:
- Mass spectrometry (MS) is crucial for protein identification and generates extensive proteomics data.
- Analyzing and storing this data efficiently presents a significant challenge in proteomics research.
Purpose of the Study:
- To design a data object model (PDOM) for protein identification data.
- To develop a parser that utilizes the PDOM to facilitate data analysis and storage.
- To create a flexible tool for managing MASCOT search results.
Main Methods:
- Developed a Protein Identification Data Object Model (PDOM).
- Created a parser compatible with HTML/XML files from MASCOT MS/MS and PMF searches.
- Implemented features for redundancy elimination and relational database output.
- Designed the system using free and open-source Java libraries.
Main Results:
- The PDOM and parser effectively process and structure protein identification data.
- Redundant information is eliminated, and data can be stored in relational databases.
- The system facilitates enhanced analysis of MASCOT search results.
- The parser is extensible for other search engines and usable as a standalone or integrated application.
Conclusions:
- The PDOM and parser offer an efficient solution for managing proteomics data from MASCOT.
- This tool aids in the storage and analysis of protein identification information.
- The extensible design serves as a template for future parser development in proteomics.
- Freely available source code promotes wider adoption and integration in bioinformatics workflows.
Related Concept Videos
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Conservation of Protein Domains
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Tagging and Fusion Proteins
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order to...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order to...

