An efficient and scalable graph modeling approach for capturing information at different levels in next generation
BMC Bioinformatics
|February 26, 2014
Summary
This study introduces an advanced graph coarsening algorithm for analyzing next-generation sequencing data. The new method efficiently models read overlaps across multiple granularity levels, improving assembly accuracy for large datasets.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Next-generation sequencing (NGS) generates vast genetic data, necessitating advanced computational tools.
- Current methods for analyzing short sequencing reads lack the complexity for efficient modeling and assembly.
- Existing approaches often use single graphs, limiting the representation of complex read relationships.
Purpose of the Study:
- To develop and evaluate an improved computational approach for analyzing and assembling next-generation sequencing data.
- To address the limitations of current methods in modeling and processing large volumes of short sequencing reads.
- To present a novel graph-theoretic algorithm for enhanced read overlap analysis.
Main Methods:
- Developed an overlap graph coarsening scheme to model read relationships at multiple levels.
- Utilized a series of graphs to represent reads and their overlaps across varying granularity.
- Integrated graph modeling and clustering for read analysis and assembly.
- Extended the algorithm to handle large simulated and real datasets from Illumina and 454 technologies.
Main Results:
- The algorithm successfully models next-generation sequencing reads at various granularity levels.
- Demonstrated efficient representation of read overlap relationships.
- Showcased scalability for large datasets.
- Validated the algorithm's practicality for both Illumina and 454 sequencing platforms.
Conclusions:
- The overlap graph theoretic algorithm provides a scalable and efficient method for next-generation sequencing data analysis.
- The multi-level graph coarsening approach enhances the modeling and clustering of sequencing reads.
- The developed method is practical and effective for diverse sequencing technologies and large-scale genomic studies.
Related Concept Videos
Next-generation Sequencing
87.9K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.9K
RNA-seq
9.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.4K
Genome Annotation and Assembly
16.7K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
16.7K
Sanger Sequencing
800.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
800.8K


