Related Experiment Video
Updated: Sep 17, 2025

12:31
In Vivo Modeling of the Morbid Human Genome using Danio rerio
Published on: August 24, 2013
20.8K
Chimeric mis-annotations of genes remain pervasive in eukaryotic non-model organisms
Andreas Bachler1, Thomas K Walsh2, Rahul V Rane3
1CSIRO, Black Mountain Laboratories, Clunies Ross Street, Canberra, ACT, 2601, Australia. Andy.Bachler@csiro.au.
BMC Genomics
|July 2, 2025
Summary
Chimeric gene mis-annotations are common in genomic datasets, particularly in non-model organisms. Machine-learning tools like Helixer can identify and correct these errors, improving genome analysis accuracy.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Accurate protein-coding gene annotation is essential for non-model organisms.
- Chimeric gene mis-annotations, where adjacent genes fuse incorrectly, are a significant problem.
- Annotation inertia propagates these errors, complicating downstream genomic analyses.
Purpose of the Study:
- To investigate the prevalence of chimeric mis-annotations in diverse genomes.
- To evaluate the effectiveness of machine-learning tools in identifying these errors.
Main Methods:
- Analysis of 30 recently annotated genomes (invertebrates, vertebrates, plants).
- Identification and confirmation of chimeric gene mis-annotations.
- Utilizing machine-learning annotation tools (e.g., Helixer) for structural prediction and splicing assessment.
Main Results:
- Identified 605 confirmed cases of chimeric mis-annotations across the studied genomes.
- The majority of errors were found in invertebrate and plant genomes.
- Machine-learning tools demonstrated efficacy in identifying these mis-annotations.
Conclusions:
- Chimeric mis-annotations are prevalent in genomic datasets, impacting data reliability.
- Machine-learning tools like Helixer offer a promising approach to refine gene models.
- Correcting these errors enhances the understanding of non-model organism genomes.
Related Concept Videos
In-vitro Mutagenesis
14.2K
To learn more about the function of a gene, researchers can observe what happens when the gene is inactivated or “knocked out,” by creating genetically engineered knockout animals. Knockout mice have been particularly useful as models for human diseases such as cancer, Parkinson’s disease, and diabetes.
14.2K
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
Exon Recombination
3.7K
The evolution of new genes is critical for speciation. Exon recombination, also known as exon shuffling or domain shuffling, is an important means of new gene formation. It is observed across vertebrates, invertebrates, and in some plants such as potatoes and sunflowers. During exon recombination, exons from the same or different genes recombine and produce new exon-intron combinations, which might evolve into new genes.
Exon shuffling follows “splice frame rules.” Each exon...
Exon shuffling follows “splice frame rules.” Each exon...
3.7K
Position-effect Variegation
6.6K
In 1928, a German botanist Emil Heitz observed the moss nuclei with a DNA binding dye. He observed that while some chromatin regions decondense and spread out in the interphase nucleus, others do not. He termed them euchromatin and heterochromatin, respectively. He proposed that the heterochromatin regions reflect a functionally inactive state of the genome. It was later confirmed that heterochromatin is transcriptionally repressed, and euchromatin is transcriptionally active chromatin.
6.6K
Cis-regulatory Sequences
10.2K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
10.2K

