Computational inference of grammars for larger-than-gene structures from annotated gene sequences
Guy Tsafnat1, Jaron Schaeffer, Andrew Clayphan
1Australian Institute of Health Innovation, University of New South Wales, Australia. guyt@unsw.edu.au
Bioinformatics (Oxford, England)
|January 25, 2011
Summary
Computational grammar inference automates the discovery of larger than gene structures (LGS), such as mobile genetic elements (MGE). A Bayesian algorithm achieved over 95% prediction accuracy for MGE structures, demonstrating robustness even with noisy data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Larger than gene structures (LGS) are DNA segments containing genes and regulatory elements, influencing organism traits like virulence and antibiotic resistance.
- Mobile genetic elements (MGE), including integrons, are significant LGS involved in horizontal gene transfer, particularly in Gram-negative bacteria.
- Expert-curated grammars represent LGS effectively for annotation and discovery but are labor-intensive and limited to known structures.
Purpose of the Study:
- To automate the discovery of LGS using computational grammar inference methods.
- To compare the efficacy of six algorithms in inferring LGS grammars from annotated DNA sequences.
- To evaluate the predictive accuracy of learned grammars against expert-defined grammars for integron gene cassette arrays.
Main Methods:
- Employed computational grammar inference algorithms to learn LGS grammars from DNA sequence data.
- Utilized annotated DNA sequences, including genes and other short sequences, as input for grammar inference.
- Compared the performance of inferred grammars against a known expert-developed grammar for specific MGE (integron gene cassette arrays).
Main Results:
- A Bayesian generalization algorithm achieved over 95% prediction of MGE structures in a large sequence corpus (F-score 75%).
- The method demonstrated robustness, maintaining a high F-score (68%) even with 100% noise in training and test datasets.
- The inferred grammars show potential for de novo prediction of LGS structures when underlying gene features are known.
Conclusions:
- Computational grammar inference provides an effective and automated approach for discovering larger than gene structures.
- The Bayesian generalization algorithm is robust and accurate for predicting MGE structures, even in the presence of data noise.
- This methodology facilitates de novo discovery of novel LGS, advancing genomic annotation and understanding of mobile DNA.
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Structure of a Gene
A gene is the fundamental unit of heredity. Every individual has two copies of each gene, one inherited from each parent. Although most people contain the same genes, there is a small fraction that is slightly different amongst people. A gene with a small difference in its sequence of DNA bases forms different alleles, contributing to different phenotypes.
However, only 1% of the DNA is composed of genes that encode proteins; the rest, 99% is non-coding DNA. This non-coding DNA performs...
However, only 1% of the DNA is composed of genes that encode proteins; the rest, 99% is non-coding DNA. This non-coding DNA performs...
Phylogenetic Trees
Phylogenetic trees come in many forms. It matters in which sequence the organisms are arranged from the bottom to the top of the tree, but the branches can rotate at their nodes without altering the information. The lines connecting individual nodes can be straight, angled, or even curved.The length of the branches can depict time or the relative amount of change among organisms. For instance, the branch length might indicate the number of amino acid changes in the sequence that underlies the...
Genome Size and the Evolution of New Genes
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
Genome Size and the Evolution of New Genes
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.


