Related Experiment Video
Updated: Dec 6, 2025

G2-seq: A High Throughput Sequencing-based Technique for Identifying Late Replicating Regions of the Genome
Published on: March 22, 2018
Finding long tandem repeats in long noisy reads
Shinichi Morishita1, Kazuki Ichikawa1, Eugene W Myers2,3
1Department of Computational Biology and Medical Sciences, Graduate School of Frontier Sciences, The University of Tokyo, Chiba 277-8562, Japan.
A new algorithm efficiently detects long tandem repeat expansions in human genomes using long-read sequencing data. This method overcomes high error rates, improving disease-associated repeat detection.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Long tandem repeat expansions (>1000 nt) are linked to diseases but are difficult to study in individual genomes due to short read lengths.
- Emerging long-read sequencing technologies (≥10,000 nt) can span these repeats but suffer from high error rates (10-20%).
- Existing algorithms for tandem repeat detection are optimized for short repeats and struggle with long reads' error rates and computational demands.
Purpose of the Study:
- To develop an efficient algorithm for detecting long tandem repeat expansions in human genomes using long-read sequencing data.
- To address the challenges posed by high error rates in long reads for repetitive element detection.
- To provide a sensitive and computationally feasible method for analyzing disease-associated repeat expansions.
Main Methods:
- Developed a novel algorithm that leverages the repetitive nature of long tandem repeats.
- Analyzed k-mer frequency distributions within fixed-size windows to estimate potential repeat regions.
- Utilized a de Bruijn graph approach to assemble k-mers into a consensus repeat unit, overcoming high error rates.
- Compared performance against the established Tandem Repeats Finder algorithm.
Main Results:
- The proposed algorithm demonstrated superior sensitivity in detecting long tandem repeat expansions compared to Tandem Repeats Finder.
- The method effectively handles the high error rates inherent in long-read sequencing data.
- Successfully identified extensive repeat regions previously challenging to analyze.
Conclusions:
- The new algorithm provides an efficient and sensitive solution for detecting long tandem repeat expansions using long-read sequencing.
- This advancement facilitates the exploration of disease associations with large repetitive elements in individual human genomes.
- The open-source implementation (mTR) is available for broader research application.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Sanger Sequencing
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
RACE - Rapid Amplification of cDNA Ends

