Related Experiment Video
Updated: Jun 26, 2025

Pattern-based Search of Epigenomic Data Using GeNemo
Published on: October 8, 2017
Indexing and searching petabase-scale nucleotide resources.
Sergey A Shiryev1, Richa Agarwala2
1Department of Health and Human Services, National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD, USA.
Pebblescout is a new tool that makes searching large nucleotide sequence databases practical for researchers. It efficiently indexes and searches massive datasets, significantly reducing analysis time and effort.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Vast nucleotide sequence data in resources like the Sequence Read Archive and GenBank are difficult to search.
- Current methods are impractical for most researchers to navigate large-scale genomic data.
Purpose of the Study:
- To introduce Pebblescout, a novel tool for indexing and searching large nucleotide sequence resources.
- To enable efficient navigation and retrieval of relevant data from massive genomic datasets.
Main Methods:
- Pebblescout employs dense sampling for indexing nucleotide sequences.
- Search functionality identifies relevant runs or assemblies based on short sequence matches.
- Results are ranked by the informativeness of the matches, providing well-defined guarantees.
Main Results:
- Eight databases indexing over 3.7 petabases were created using Pebblescout.
- The tool effectively finds relevant subsets of large nucleotide resources across various query lengths.
- Pebblescout demonstrates favorable performance compared to existing tools like MetaGraph and Sourmash.
Conclusions:
- Pebblescout significantly reduces the effort required for downstream analysis of large nucleotide resources.
- The tool offers a data-driven approach for efficient exploration of massive genomic data.
- Pebblescout provides a practical solution for researchers dealing with rapidly growing sequence databases.
More Related Videos
10:41Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
Published on: May 9, 2017
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
Related Concept Videos
Nucleic Acids and Nucleotides
Deoxyribonucleic Acid (DNA)
DNA is the genetic material in all living organisms, ranging from single-celled bacteria to multicellular mammals. It is in the nucleus of eukaryotes and the organelles such as chloroplasts and mitochondria....
Protein Families
Sanger Sequencing