Related Experiment Videos
Seven clusters in genomic triplet distributions
Alexander N Gorban1, Andrei Y Zinovyev, Tatyana G Popova
1Institute of Computational Modeling, Russian Academy of Science.
In Silico Biology
|February 18, 2004
Summary
Unsupervised gene detection is possible due to cluster structures in genome triplet frequencies. This study visualizes these clusters, achieving over 90% accuracy in identifying protein-coding regions without prior gene data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Recent advancements propose unsupervised gene-detection algorithms, eliminating the need for pre-existing gene datasets.
- The feasibility of unsupervised gene detection is linked to inherent cluster structures within oligomer frequency distributions.
Purpose of the Study:
- To investigate the cluster structure within genomic triplet frequencies.
- To explore the potential of data visualization for analyzing genomic sequence properties.
Main Methods:
- Analysis of complete genomic sequences using a pure data exploration strategy.
- Visualization of triplet frequency tables within a sliding window.
- Examination of 64-dimensional vectors representing triplet frequencies.
Main Results:
- A well-detectable cluster structure was identified in the distribution of triplet frequencies.
- The structure comprises seven distinct clusters.
- These clusters accurately correspond to protein-coding regions (in three possible phases on complementary strands) and non-coding regions, with nucleotide-level accuracy exceeding 90%.
Conclusions:
- Visualizing and understanding genomic cluster structures aids in analyzing gene-prediction tool performance.
- This method, not requiring Open Reading Frame (ORF) extraction, is applicable to unassembled genomes.