Related Experiment Video
Updated: Aug 7, 2026

Genomic MRI - a Public Resource for Studying Sequence Patterns within Genomic DNA
Published on: May 9, 2011
Statistical measures of the structure of genomic sequences: entropy, complexity, and position information
Yuriy L Orlov1, Rene Te Boekhorst, Irina I Abnizova
1Institute of Cytology and Genetics SB RAS, Lavrentieva Ave., 10, Novosibirsk 630090, Russia. orlov@bionet.nsc.ru
Abstract:
Identifying regions of DNA with extreme statistical characteristics is an important aspect of the structural analysis of complete genomes. Linguistic methods, mainly based on estimating word frequency, can be used for this as they allow for the delineation of regions of low complexity. Low complexity may be due to biased nucleotide composition, by tandem- or dispersed repeats, by palindrome-hairpin structures, as well as by a combination of all these features. We developed software tools in which various numerical measures of text complexity are implemented, including combinatorial and linguistic ones. We also added Hurst exponent estimate to the software to measure dependencies in DNA sequences. By applying these tools to various functional genomic regions, we demonstrate that the complexity of introns and regulatory regions is lower than that of coding regions, whilst Hurst exponent is larger. Further analysis of promoter sequences revealed that the lower complexity of these regions is associated with long-range correlations caused by transcription factor binding sites.
Related Concept Videos
Organization of Genes
Organization of Genes
Modern Molecular Taxonomy
Evolutionary Relationships through Genome Comparisons
Entropy
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...

