Related Experiment Video
Updated: Jun 10, 2026

High-throughput Identification of Gene Regulatory Sequences Using Next-generation Sequencing of Circular Chromosome Conformation Capture (4C-seq)
Published on: October 5, 2018
Graph-based clustering and characterization of repetitive sequences in next-generation sequencing data
Petr Novák1, Pavel Neumann, Jirí Macas
1Biology Centre ASCR, Institute of Plant Molecular Biology, Branisovska 31, Ceske Budejovice, CZ-37005, Czech Republic.
This study presents a novel graph-based approach for analyzing plant repetitive DNA sequences using 454 sequencing data. This method efficiently characterizes repeat families and their variability in plant genomes.
Area of Science:
- Genomics
- Bioinformatics
- Plant Science
Background:
- Repetitive sequences constitute the majority of higher plant nuclear DNA, crucial for understanding genome structure and evolution.
- Genome-wide characterization of these abundant and diverse repetitive elements is challenging.
- Massively-parallel sequencing, particularly 454 sequencing, offers a promising avenue for efficient repeat analysis.
Purpose of the Study:
- To adapt a graph-based approach for partitioning whole-genome 454 sequence reads into clusters representing individual repeat families.
- To develop and utilize data mining tools for efficient repeat characterization in plant genomes.
- To assess the proportion, composition, and variability of repeats within plant genomes.
Main Methods:
- Utilized a graph-based approach for similarity-based partitioning of 454 sequence reads.
- Applied cluster size information to estimate repeat proportions and composition in *Pisum sativum* and *Glycine max*.
- Employed statistical analysis and visual inspection of cluster graph topology with the SeqGrapheR program.
Main Results:
- Successfully clustered 454 sequence reads to represent individual repeat families.
- Quantified repeat proportions and composition in *Pisum sativum* and *Glycine max* genomes.
- SeqGrapheR facilitated the distinction of basic repeat types and investigation of sequence variability within families.
Conclusions:
- Graph-based analysis provides an efficient method for characterizing repetitive regions in plant genomes.
- The graph representation aids in assessing repeat family variability, evolutionary divergence, and discovery of novel elements.
- This approach supports the subsequent assembly of consensus sequences for identified repeats.
Related Concept Videos
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Evolutionary Relationships through Genome Comparisons
Genome Annotation and Assembly
Modern Molecular Taxonomy

