Related Experiment Video
Updated: Dec 26, 2025

De novo Identification of Actively Translated Open Reading Frames with Ribosome Profiling Data
Published on: February 18, 2022
Efficient Construction of a Complete Index for Pan-Genomics Read Alignment.
Alan Kuhnle1,2, Taher Mun3, Christina Boucher2
1Department of Computer Science, Florida State University, Tallahassee, Florida.
This study introduces an efficient method for building FM-indexes for large genomic databases, improving memory and time efficiency for indexing and querying human genomes.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Short-read aligners commonly use FM-indexes, which struggle to scale for large genomic databases.
- Existing FM-index components, the rank data structure and suffix array (SA) sample, present scalability challenges.
- Run-length compression can reduce the rank data structure size, but efficiently managing the SA sample remained an issue.
Purpose of the Study:
- To develop an efficient method for constructing FM-indexes for large-scale genomic databases.
- To address the challenge of efficiently building the suffix array (SA) sample for FM-indexes.
- To improve the performance of indexing and querying large collections of genomes.
Main Methods:
- Developed a novel approach for efficiently constructing the suffix array (SA) sample for FM-indexes.
- Compared the proposed SA sample construction method against state-of-the-art techniques.
- Applied the developed method to index partial and whole human genomes.
Main Results:
- The new method demonstrates superior speed and space efficiency for constructing SA samples, especially on repetitive genomic databases.
- Indexing human genomes using the new method shows improvements in both memory usage and time compared to Bowtie.
- The approach also outperforms the CHIC method in terms of query time and indexing memory requirements.
Conclusions:
- The developed method provides an efficient solution for building FM-indexes for large genomic databases.
- This advancement enables faster and more memory-efficient indexing and querying of extensive genomic collections.
- The findings have significant implications for large-scale genomic data analysis and bioinformatics tools.
More Related Videos
12:08Hybrid De Novo Genome Assembly for the Generation of Complete Genomes of Urinary Bacteria using Short- and Long-read Sequencing Technologies
Published on: August 20, 2021
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018