Fast characterization of segmental duplications in genome assemblies
Ibrahim Numanagic1,2, Alim S Gökkaya3, Lillian Zhang1
1Computer Science and Artificial Intelligence Laboratory, Cambridge, MA, USA.
Bioinformatics (Oxford, England)
|November 14, 2018
Summary
Segmental duplications (SDs) are crucial for evolution but challenge genome assembly. A new tool, SEDEF, rapidly and accurately detects these complex DNA regions, improving genomic analysis.
Area of Science:
- Genomics
- Bioinformatics
- Evolutionary Biology
Background:
- Segmental duplications (SDs) are large, highly similar DNA segments that contribute to genome evolution and structural variation.
- SDs are a major obstacle in de novo genome assembly, leading to misassemblies, collapsed regions, or missing data.
- Accurate characterization of SDs is essential for understanding genome architecture and evolution, yet existing tools are cumbersome.
Purpose of the Study:
- To develop a rapid, accurate, and user-friendly tool for detecting segmental duplications within genome assemblies.
- To address the limitations of existing methods for analyzing SDs in the context of genome assembly.
Main Methods:
- Introduced the SEgmental Duplication Evaluation Framework (SEDEF).
- Employed sophisticated filtering strategies, including Jaccard similarity and local chaining, for SD detection.
- Focused on capturing a wider range of pairwise errors between segments (up to 25%) compared to previous methods.
Main Results:
- SEDEF achieves substantial speed-up over existing tools like Whole-Genome Assembly Comparison (WGAC), reducing runtimes from weeks to minutes.
- SEDEF accurately detects SDs, providing a more comprehensive analysis of genomic structure.
- The framework enables deeper tracking of genome evolutionary history by accounting for a greater degree of segment similarity.
Conclusions:
- SEDEF offers a significant advancement in the rapid and accurate detection of segmental duplications in genome assemblies.
- The tool overcomes previous limitations, facilitating better understanding of genome evolution and structural variation.
- SEDEF is publicly available, promoting wider adoption and research in genomics.
Related Concept Videos
Genome Annotation and Assembly
21.0K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
21.0K
Genome Size and the Evolution of New Genes
9.1K
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
9.1K
Centrosome Duplication
5.0K
The primary microtubule organizing center (MTOC) in animal cells is the centrosome. A centrosome has two cylindrical centrioles at its core. Each centriole consists of nine sets of three microtubules held together by proteins. The centrioles are positioned at right angles to each other and surrounded by a shapeless protein cloud called the pericentriolar matrix, or pericentriolar material (PCM).
To ensure that each daughter cell receives a centrosome after cell division, centrosome duplication...
To ensure that each daughter cell receives a centrosome after cell division, centrosome duplication...
5.0K
Gene Duplication and Divergence
8.0K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
8.0K
Gene Evolution - Fast or Slow?
8.2K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
8.2K
Duplication of Chromatin Structure
7.4K
The process of chromosome duplication during cell division requires genome-wide disruption and re-assembly of chromatin. The chromatin structure must be accurately inherited, reassembled, and maintained in the daughter cells to ensure lineage propagation.
The basic unit of the chromatin is the nucleosome, consisting of DNA wrapped around octameric histone proteins and short stretches of linker DNA separating individual nucleosomes. The histone proteins within the nucleosome have their...
The basic unit of the chromatin is the nucleosome, consisting of DNA wrapped around octameric histone proteins and short stretches of linker DNA separating individual nucleosomes. The histone proteins within the nucleosome have their...
7.4K


