Related Experiment Video
Updated: May 26, 2026

Electrophoretic Analysis of Replication Through Structure-Prone DNA Repeats Within the SV40-Based Human Episome
Published on: September 13, 2024
Repetitive elements may comprise over two-thirds of the human genome
A P Jason de Koning1, Wanjun Gu, Todd A Castoe
1Department of Biochemistry and Molecular Genetics, School of Medicine, University of Colorado, Aurora, Colorado, USA.
Abstract:
Transposable elements (TEs) are conventionally identified in eukaryotic genomes by alignment to consensus element sequences. Using this approach, about half of the human genome has been previously identified as TEs and low-complexity repeats. We recently developed a highly sensitive alternative de novo strategy, P-clouds, that instead searches for clusters of high-abundance oligonucleotides that are related in sequence space (oligo "clouds"). We show here that P-clouds predicts >840 Mbp of additional repetitive sequences in the human genome, thus suggesting that 66%-69% of the human genome is repetitive or repeat-derived. To investigate this remarkable difference, we conducted detailed analyses of the ability of both P-clouds and a commonly used conventional approach, RepeatMasker (RM), to detect different sized fragments of the highly abundant human Alu and MIR SINEs. RM can have surprisingly low sensitivity for even moderately long fragments, in contrast to P-clouds, which has good sensitivity down to small fragment sizes (∼25 bp). Although short fragments have a high intrinsic probability of being false positives, we performed a probabilistic annotation that reflects this fact. We further developed "element-specific" P-clouds (ESPs) to identify novel Alu and MIR SINE elements, and using it we identified ∼100 Mb of previously unannotated human elements. ESP estimates of new MIR sequences are in good agreement with RM-based predictions of the amount that RM missed. These results highlight the need for combined, probabilistic genome annotation approaches and suggest that the human genome consists of substantially more repetitive sequence than previously believed.
Related Concept Videos
Organization of Genes
Chromosome Structure
The centromere is a DNA sequence that links sister chromatids. This is also where kinetochores, protein complexes to which spindle microtubules attach, are constructed after the chromosome is replicated. The kinetochores allow the spindle microtubules to move the chromosomes within the cell during cell division.
Telomeres consist of non-coding repetitive nucleotide...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Genomic DNA in Eukaryotes
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are characterized.
Duplication of Chromatin Structure
The basic unit of the chromatin is the nucleosome, consisting of DNA wrapped around octameric histone proteins and short stretches of linker DNA separating individual nucleosomes. The histone proteins within the nucleosome have their...

