Related Experiment Video
Updated: Mar 22, 2026

An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Identification and Correction of Erroneous Protein Sequences in Public Databases
1Institute of Enzymology, Research Centre for Natural Sciences, Hungarian Academy of Sciences, 286, Budapest, H-1519, Hungary. patthy.laszlo@ttk.mta.hu.
Predicting protein-coding gene structures is challenging, leading to errors in public databases. New MisPred and FixPred methods identify and correct these erroneous sequences by checking for conflicting protein features.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Accurate prediction of protein-coding gene structures in higher eukaryotes is complex.
- Public sequence databases are often contaminated with mispredicted, erroneous sequences.
- Mispredictions significantly impact the reliability of genome-scale sequence analyses.
Purpose of the Study:
- To introduce novel computational approaches for identifying and correcting erroneous protein sequences.
- To improve the quality and reliability of sequence data in public databases.
Main Methods:
- Development of the MisPred approach for identifying erroneous sequences.
- Development of the FixPred approach for correcting erroneous sequences.
- Utilizing the principle that erroneous protein sequences exhibit feature conflicts with known protein characteristics.
Main Results:
- MisPred and FixPred provide a framework for detecting and rectifying sequence errors.
- These methods leverage existing knowledge of protein features to flag anomalies.
- The approaches aim to reduce the contamination of public sequence databases.
Conclusions:
- The MisPred and FixPred approaches offer valuable tools for enhancing the accuracy of protein sequence data.
- Improving sequence data quality is crucial for robust genomic research.
- Addressing sequence mispredictions supports more reliable biological interpretations from large-scale analyses.
More Related Videos
07:49Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
Published on: August 16, 2017
07:38Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
Published on: April 11, 2019
Related Concept Videos
Protein Families
Mismatch Repair
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Conservation of Protein Domains
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...