Related Experiment Video
Updated: Jun 21, 2026

09:40
Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
Parametric complexity of sequence assembly: theory and applications to next generation sequencing.
Niranjan Nagarajan1, Mihai Pop
1Center for Bioinformatics and Computational Biology, Institute for Advanced Computer Studies, University of Maryland, College Park, Maryland 20742, USA.
Summary
New DNA sequencing generates vast data, but genome assembly remains challenging. This study theoretically analyzes assembly complexity, linking repeats, read length, and coverage to improve algorithms.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Advancements in DNA sequencing technologies have led to an unprecedented volume of genomic data.
- Genome assembly, the process of reconstructing genomes from short DNA sequences, remains a significant algorithmic challenge.
- Current genome assembly methods primarily rely on heuristic solutions, lacking theoretical underpinnings for complex data.
Purpose of the Study:
- To provide a theoretical framework for understanding the parametric complexity of genome assembly.
- To identify key factors influencing the difficulty of genome assembly problems.
- To guide the development of more robust and efficient genome assembly algorithms.
Main Methods:
- Theoretical analysis of genome assembly.
- Exploration of the relationship between repeat complexity, read lengths, overlap lengths, and coverage.
- Identification of parameters that define computationally 'hard' instances of the assembly problem.
Main Results:
- Established connections between repeat complexity, read length, overlap length, and coverage in determining assembly difficulty.
- Identified specific parameters that contribute to challenging genome assembly instances.
- Provided a theoretical basis for understanding the limitations of current heuristic assembly approaches.
Conclusions:
- A deeper theoretical understanding of genome assembly complexity is crucial for advancing the field.
- The study suggests avenues for rigorously extending existing genome assemblers.
- Future theoretical investigations are needed to further refine assembly algorithms and address new sequencing technologies.
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Maxam-Gilbert Sequencing
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Sanger Sequencing
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...

