Efficiency of Learned Indexes on Genome Spectra
Md Hasin Abrar1, Paul Medvedev2, Giorgio Vinciguerra3
1Department of Computer Science and Engineering, Penn State, University Park, USA.
A new measure, Canonical Piecewise Linear approximability (CaPLa), accurately predicts the performance of data structures for genomic k-mers. This measure reveals wide variations in k-mer patterns across life, improving bioinformatics tool efficiency.
Area of Science:
- Bioinformatics
- Computational Biology
- Data Structures
Background:
- Bioinformatic tools rely on data structures for genomic k-mers.
- Efficient data structures must leverage inherent data patterns.
- Learned indexes offer efficient k-mer multiset rank approximation but lack predictive performance analysis.
Purpose of the Study:
- Develop a novel measure for piecewise-linear approximability of data.
- Introduce Canonical Piecewise Linear approximability (CaPLa) to predict data structure performance.
- Analyze the variability of genomic k-mer patterns.
Main Methods:
- Developed CaPLa based on power-law model deviations.
- Proved CaPLa properties and created an efficient computation algorithm.
- Applied CaPLa to predict space bounds for real-world data structures.
Main Results:
- CaPLa accurately predicts space bounds for data structures on genomic data.
- Empirical analysis of over 500 genomes shows significant CaPLa variation across the tree of life and within genomes.
- Identified factors contributing to genomic k-mer multisets differing from random ones.
Conclusions:
- CaPLa is a robust measure for predicting data structure performance on genomic k-mers.
- Genomic k-mer patterns exhibit substantial variability, impacting data structure efficiency.
- CaPLa provides insights into optimizing bioinformatics tools for diverse genomic datasets.
More Related Videos
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
07:11Digital PCR-based Competitive Index for High-throughput Analysis of Fitness in Salmonella
Published on: May 13, 2019
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
DNA Microarrays
Gene Evolution - Fast or Slow?
In contrast, regions which code...
Expected Frequencies in Goodness-of-Fit Tests
Modern Molecular Taxonomy
Genome Size and the Evolution of New Genes
