Related Experiment Video
Updated: Sep 21, 2025

Metagenomic Analysis of Silage
Published on: January 13, 2017
Use of 6 Nucleotide Length Words to Study the Complexity of Gene Sequences from Different Organisms
Eugene Korotkov1, Konstantin Zaytsev2, Alexey Fedorov2
1Institute of Bioengineering, Federal Research Center of Biotechnology of the Russian Academy of Sciences, 119071 Moscow, Russia.
Abstract:
In this paper, we attempted to find a relation between bacteria living conditions and their genome algorithmic complexity. We developed a probabilistic mathematical method for the evaluation of k-words (6 bases length) occurrence irregularity in bacterial gene coding sequences. For this, the coding sequences from different bacterial genomes were analyzed and as an index of k-words occurrence irregularity, we used W, which has a distribution similar to normal. The research results for bacterial genomes show that they can be divided into two uneven groups. First, the smaller one has W in the interval from 170 to 475, while for the second it is from 475 to 875. Plants, metazoan and virus genomes also have W in the same interval as the first bacterial group. We suggested that second bacterial group coding sequences are much less susceptible to evolutionary changes than the first group ones. It is also discussed to use the W index as a biological stress value.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome Size and the Evolution of New Genes
Genomics
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
DNA as a Genetic Template
Modern Molecular Taxonomy

