Related Experiment Video
Updated: Aug 17, 2026

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
Published on: March 22, 2018
De novo identification of repeat families in large genomes
Alkes L Price1, Neil C Jones, Pavel A Pevzner
1Department of Computer Science and Engineering, University of California San Diego La Jolla, CA 92093-0114, USA.
Motivation:
De novo repeat family identification is a challenging algorithmic problem of great practical importance. As the number of genome sequencing projects increases, there is a pressing need to identify the repeat families present in large, newly sequenced genomes. We develop a new method for de novo identification of repeat families via extension of consensus seeds; our method enables a rigorous definition of repeat boundaries, a key issue in repeat analysis.
Results:
Our RepeatScout algorithm is more sensitive and is orders of magnitude faster than RECON, the dominant tool for de novo repeat family identification in newly sequenced genomes. Using RepeatScout, we estimate that approximately 2% of the human genome and 4% of mouse and rat genomes consist of previously unannotated repetitive sequence.
Availability:
Source code is available for download at http://www-cse.ucsd.edu/groups/bioinformatics/software.html
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are characterized.
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Modern Molecular Taxonomy
Genome Annotation and Assembly

