Related Experiment Video
Updated: May 4, 2026

08:43
Metagenomic Analysis of Silage
Published on: January 13, 2017
18.7K
A bioinformatician's guide to the forefront of suffix array construction algorithms
Anish Man Singh Shrestha1, Martin C Frith, Paul Horton
1Computational Biology Research Center, AIST, Tokyo, Japan. computome@gmail.com.
Briefings in Bioinformatics
|January 14, 2014
Summary
This study explains the SA-IS algorithm for suffix array construction and introduces DisLex for modified suffix arrays. These methods aid inexact matching in bioinformatics, supporting spaced and subset seeds in biological applications.
Area of Science:
- Bioinformatics
- Computational Biology
- Data Structures
Background:
- Suffix arrays are crucial text-indexing data structures in bioinformatics.
- Efficient construction of suffix arrays is essential for various biological applications.
Purpose of the Study:
- To provide an accessible explanation of the state-of-the-art SA-IS algorithm for suffix array construction.
- To introduce DisLex, a novel technique for creating modified suffix arrays.
- To enable simple inexact matching for biological sequence analysis.
Main Methods:
- Detailed exposition of the SA-IS (Suffix Array - Induced Sorting) algorithm.
- Description of the DisLex technique for modifying suffix arrays.
- Application of modified suffix arrays for spaced and subset seed matching.
Main Results:
- The SA-IS algorithm is presented as a state-of-the-art method for suffix array construction.
- DisLex enables standard algorithms to generate modified suffix arrays.
- The modified suffix arrays facilitate inexact matching crucial for biological applications.
Conclusions:
- The SA-IS algorithm and DisLex technique offer valuable tools for bioinformatics.
- These advancements support the use of spaced and subset seeds in biological data analysis.
- Accessible explanations and new techniques enhance the utility of suffix arrays in the field.
Related Concept Videos
Genome Annotation and Assembly
16.7K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
16.7K
Sanger Sequencing
800.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
800.8K
Modern Molecular Taxonomy
836
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
836
Next-generation Sequencing
87.9K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.9K
RNA-seq
9.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.4K
DNA Microarrays
16.8K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
16.8K

