Related Experiment Video
Updated: Aug 7, 2025

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.3K
Coverage-preserving sparsification of overlap graphs for long-read assembly
1Department of Computational and Data Sciences, Indian Institute of Science, Bengaluru, Karnataka 560012, India.
Bioinformatics (Oxford, England)
|March 9, 2023
Summary
Genome assembly graph models are crucial for de novo assembly. We show the standard string graph model can create coverage gaps, but a new method retains key reads to fix this.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Read-overlap graphs are essential for de novo genome assembly.
- String graph models are widely used for graph sparsification to improve assembly contiguity.
- Coverage preservation is critical, especially for complex genomes, to avoid losing haplotype information.
Purpose of the Study:
- To develop a theoretical framework for analyzing coverage-preserving properties of graph models in genome assembly.
- To identify limitations of the standard string graph model regarding coverage preservation.
- To propose practical heuristics for retaining essential contained reads and mitigating coverage gaps.
Main Methods:
- Developed a novel theoretical framework to analyze coverage-preserving properties of graph models.
- Proved de Bruijn and overlap graphs are coverage-preserving.
- Demonstrated the standard string graph model's deficiency and identified contained read removal as a cause of coverage gaps.
- Proposed and validated heuristics for retaining a small fraction of contained reads.
Main Results:
- The standard string graph model is not guaranteed to be coverage-preserving.
- Ignoring contained reads can introduce significant coverage gaps (e.g., 50 gaps on average for HG002 nanopore data).
- Proposed heuristics effectively close most coverage gaps by retaining only 1-2% of contained reads.
Conclusions:
- The standard string graph model requires modifications to ensure coverage preservation.
- The proposed heuristics offer a practical solution to prevent coverage gaps in genome assembly.
- This work enhances the reliability of de novo genome assembly for complex genomic datasets.
More Related Videos
Related Concept Videos
Genome Annotation and Assembly
19.0K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.0K
Long-patch Base Excision Repair
7.1K
Since the discovery of the two BER pathways, there has been a debate about how a cell chooses one pathway over the other and the factors determining this selection. Numerous in vitro experiments have pointed out multiple determinants for the sub-pathway selection. These are:
7.1K
Multi-species Conserved Sequences
4.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.0K
Sanger Sequencing
755.5K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
755.5K
Gene Duplication and Divergence
6.2K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
6.2K
Genome Copying Errors
4.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.3K

