Related Experiment Video
Updated: Feb 3, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.6K
Quantifying local randomness in human DNA and RNA sequences using Erdös motifs
Wentian Li1, Dimitrios Thanos2, Astero Provata3
1The Robert S. Boas Center for Genomics and Human Genetics, The Feinstein Institute for Medical Research, Northwell Health, Manhasset, NY, USA.
Journal of Theoretical Biology
|October 19, 2018
Summary
Researchers investigated Erdös sequences in human DNA, finding purine/pyrimidine motifs are underrepresented while strong/weak motifs are overrepresented. These findings shed light on genomic sequence patterns and their potential implications.
Area of Science:
- Genomics
- Bioinformatics
- Number Theory
Background:
- Paul Erdös posed a question in 1932 regarding random walks and binary sequences achieving minimal discrepancy.
- Constructing finite sequences with low discrepancy properties, termed "Erdös sequences," has been an area of interest.
- Erdös sequences exhibit local randomness, characterized by balanced short subsequence frequencies and avoidance of periodic patterns.
Purpose of the Study:
- To analyze the frequency of specific Erdös motifs in human DNA using various nucleotide-to-binary mappings.
- To investigate the representation of purine/pyrimidine and strong/weak nucleotide-based motifs within the human genome and transcriptome.
Main Methods:
- Utilized three distinct nucleotides-to-binary mappings to analyze Erdös motifs of length-10.
- Examined the frequency and density of these motifs in human genomic DNA and messenger RNA sequences.
- Applied concepts from number theory, specifically related to discrepancy and sequences like Thue-Morse.
Main Results:
- Purine (A, G) / pyrimidine (C, T) based Erdös motifs are significantly underrepresented in the human genome.
- Strong (G, C) / weak (A, T) based Erdös motifs are slightly overrepresented in the human genome.
- Erdös motifs derived from all three mappings combined are slightly underrepresented, while strong/weak motifs are greatly overrepresented in human messenger RNA.
Conclusions:
- Human DNA exhibits specific underrepresentation of purine/pyrimidine Erdös motifs and overrepresentation of strong/weak motifs.
- The observed patterns suggest a non-random distribution of these motifs within the human genome and transcriptome.
- These findings contribute to understanding the structural and potentially functional characteristics of genomic sequences.
Related Concept Videos
DNA-only Transposons
17.5K
DNA-only transposons are called autonomous transposons since they code for the enzyme transposase that is required for the transposition mechanism. Insertion of transposons can alter gene functions in multiple ways. They can mutate the gene, alter gene expression by introducing a novel promoter or insulator sequence, introduce new splice sites, and change the mRNA transcripts produced, or remodel chromatin structure.
The donor site from where the transposon is excised is either degraded or...
The donor site from where the transposon is excised is either degraded or...
17.5K
Bacterial RNA Polymerase
32.8K
Unlike eukaryotes, bacteria use a single RNA Polymerase (RNAP) to transcribe all genes. The different subunits of bacterial RNAPhave distinct functions. The multisubunit structure of the bacterial RNAP helps the enzyme to maintain catalytic function, facilitate assembly, interact with DNA and RNA, and self-regulate its activity.
In most genes, the transcription site is a single base present upstream of the coding sequence. Though RNAP is a catalytically efficient enzyme, it does not recognize...
In most genes, the transcription site is a single base present upstream of the coding sequence. Though RNAP is a catalytically efficient enzyme, it does not recognize...
32.8K
RNA Splicing
60.6K
Splicing is the process by which eukaryotic RNA is edited before its translation into protein. The RNA strand transcribed from eukaryotic DNA is called the primary transcript. The primary transcripts that become mRNAs are called precursor messenger RNAs (pre-mRNAs). Eukaryotic pre-mRNA contains alternating sequences of exons and introns. Exons are nucleotide sequences that code for proteins, whereas introns are the non-coding regions. In RNA splicing, introns are removed and exons are bonded...
60.6K
RNA Stability
35.7K
Intact DNA strands can be found in fossils, while scientists sometimes struggle to keep RNA intact under laboratory conditions. The structural variations between RNA and DNA underlie the differences in their stability and longevity. Because DNA is double-stranded, it is inherently more stable. The single-stranded structure of RNA is less stable but also more flexible and can form weak internal bonds. Additionally, most RNAs in the cell are relatively short, while DNA can be up to 250 million...
35.7K
Eukaryotic RNA Polymerases
27.1K
RNA Polymerase (RNAP) is conserved in all animals, with bacterial, archaeal, and eukaryotic RNAPs sharing significant sequence, structural, and functional similarities. Among the three eukaryotic RNAPs, RNA Polymerase II is most similar to bacterial RNAP in terms of both structural organization and folding topologies of the enzyme subunits. However, these similarities are not reflected in their mechanism of action.
All three eukaryotic RNAPs require specific transcription factors, of which the...
All three eukaryotic RNAPs require specific transcription factors, of which the...
27.1K
RNA Editing
9.9K
RNA editing is a post-transcriptional modification where a precursor mRNA (pre-mRNA) nucleotide sequence is changed by base insertion, deletion, or modification. The extent of RNA editing varies from a few hundred bases, in mitochondrial DNA of trypanosomes, to a just single base, in nuclear genes of mammals. Even a single base change in the pre-mRNA can convert a codon for one amino acid into the codon for another amino acid or a stop codon. This type of re-coding can significantly affect the...
9.9K

