Related Experiment Video
Updated: Sep 25, 2025

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
KmerKeys: a web resource for searching indexed genome assemblies and variants
Dmitri S Pavlichin1, HoJoon Lee1, Stephanie U Greer1
1Division of Oncology, Department of Medicine, Stanford University School of Medicine, Stanford, CA, 94305, USA.
KmerKeys is a new web application that accelerates genome sequence analysis by efficiently indexing and searching massive k-mer datasets. It offers fast, scalable, and flexible querying for genomic data, including variant information.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- K-mers (short DNA sequences) are crucial for genome analysis, but their large scale presents computational challenges.
- Analyzing billions of k-mers in human genome assemblies requires significant computational resources.
Purpose of the Study:
- To develop a performant and scalable solution for analyzing large-scale k-mer data in genome assemblies.
- To introduce KmerKeys, a web application enabling rapid and flexible k-mer searching.
Main Methods:
- Developed a novel indexing data structure using a hash table optimized for short sequence keys.
- Implemented cache-friendly hash tables, memory mapping, and massive parallel processing for speed and efficiency.
- Enabled both exact and fuzzy sequence searches within genome assemblies.
Main Results:
- KmerKeys provides rapid query speeds for cloud-based genome assembly analysis.
- The system efficiently handles and searches large collections of human genome assembly information.
- Allows integration of variant databases and metadata, such as gnomAD, for enhanced analysis.
Conclusions:
- KmerKeys offers a scalable and efficient solution to the computational challenges of k-mer analysis in genomics.
- The application facilitates advanced genomic data analysis by enabling flexible and fast querying.
- KmerKeys supports the incorporation of diverse genomic information, paving the way for future sequencing analysis advancements.
More Related Videos
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
Related Concept Videos
Genome Annotation and Assembly
Evolutionary Relationships through Genome Comparisons
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...