Related Experiment Video
Updated: Dec 1, 2025

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
3.8K
Understanding the causes of errors in eukaryotic protein-coding gene prediction: a case study of primate proteomes
Corentin Meyer1, Nicolas Scalzitti1, Anne Jeannin-Girardon1
1Department of Computer Science, ICube, CNRS, University of Strasbourg, Strasbourg, France.
BMC Bioinformatics
|November 11, 2020
Summary
Gene prediction errors are common in primate proteomes, impacting up to 50% of sequences. This study identifies causes and proposes methods to improve protein sequence accuracy.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Accurate genome annotation is crucial but challenging, especially for eukaryotic protein-coding genes due to complex exon-intron structures.
- Existing gene prediction algorithms can produce significant errors affecting downstream analyses.
Purpose of the Study:
- To investigate the prevalence and causes of gene prediction errors in primate proteomes.
- To develop methods for improving protein sequence quality.
Main Methods:
- Analyzed 176,478 proteins from ten primate proteomes, using human proteins as a reference.
- Characterized gene prediction errors, focusing on mismatched segments.
- Identified potential causes for mispredictions and developed a proof-of-concept improvement method.
Main Results:
- Detected 82,305 potential errors, including deletions, insertions, and mismatched segments.
- Identified causes for approximately half of the mismatched sequence errors.
- Successfully proposed improved sequences for 603 primate proteins.
Conclusions:
- Gene prediction errors affect up to 50% of primate protein sequences.
- Causes include undetermined genome regions, sequencing issues, and model limitations.
- Existing genome data can be leveraged to enhance protein sequence quality.
More Related Videos
Related Concept Videos
Leaky Scanning
5.5K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.5K
Improving Translational Accuracy
12.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
12.9K
Synteny and Evolution
3.6K
John H. Renwick first coined the term “synteny” in 1971, which refers to the genes present on the same chromosomes, even if they are not genetically linked. The species with common ancestry tend to show conserved syntenic regions. Therefore, the concept of synteny is nowadays used to describe the evolutionary relationship between species.
Around 80 million years ago, the human and mice lineages diverged from the common ancestor. During the course of evolution, the ancestral...
Around 80 million years ago, the human and mice lineages diverged from the common ancestor. During the course of evolution, the ancestral...
3.6K
Genome Copying Errors
4.8K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.8K
Conservation of Protein Domains Over Different Proteins
13.7K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
13.7K
Multi-species Conserved Sequences
4.5K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.5K

