Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multi-species Conserved Sequences02:51

Multi-species Conserved Sequences

4.5K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale  studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.5K
RNA-seq03:21

RNA-seq

11.5K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.5K
Per-Unit Sequence Models01:26

Per-Unit Sequence Models

351
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
351
Mismatch Repair01:20

Mismatch Repair

6.1K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
6.1K
Sequences01:29

Sequences

141
Sequences are fundamental mathematical objects consisting of ordered lists of numbers that follow a specific rule or pattern. Sequences are critical in various mathematical concepts, including calculus, series, and number theory. They can model real-world phenomena such as population growth, financial investments, and physical processes like the diminishing height of a bouncing ball.Each number in a sequence is referred to as a term. Typically, the terms are denoted as a1, a2, a3,…, where...
141

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Genome misassembly detection using Stash: A data structure based on stochastic tile hashing.

PloS one·2026
Same author

Inviting Participants to the Table: Application of the Data Placemats for Disseminating Research Results.

Progress in community health partnerships : research, education, and action·2026
Same author

Palliative care interventions for thoracic cancer: a systematic review and meta-analysis identifying the core elements.

European respiratory review : an official journal of the European Respiratory Society·2026
Same author

Advances in Diagnosis-Pulmonology.

Journal of thoracic oncology : official publication of the International Association for the Study of Lung Cancer·2026
Same author

Country-level determinants of healthcare access for refugee populations in a host country: a comparative case study.

International journal for equity in health·2026
Same author

AIEdit: Alignment-free genome assembly polisher trained on spaced seed match patterns.

PLoS computational biology·2026

Related Experiment Video

Updated: Dec 15, 2025

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.6K

Mismatch-tolerant, alignment-free sequence classification using multiple spaced seeds and multiindex Bloom filters.

Justin Chu1, Hamid Mohamadi2, Emre Erhan2

  • 1Canada's Michael Smith Genome Sciences Centre, BC Cancer, Vancouver, BC V5Z 4S6, Canada; cjustin@bcgsc.ca ibirol@bcgsc.ca.

Proceedings of the National Academy of Sciences of the United States of America
|July 10, 2020
PubMed
Summary

A new multi-index Bloom Filter (miBF) improves bioinformatics analysis by efficiently classifying sequencing data. This method enhances sensitivity and specificity for read-binning and taxonomic assignment while reducing computational costs.

Keywords:
Bloom filtersalignment-freeprobabilistic data structuressequence classificationspaced seeds

More Related Videos

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

7.4K
A Practical Guide to Phylogenetics for Nonexperts
12:00

A Practical Guide to Phylogenetics for Nonexperts

Published on: February 5, 2014

35.9K

Related Experiment Videos

Last Updated: Dec 15, 2025

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.6K
Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group
07:49

Creating and Applying a Reference to Facilitate the Discussion and Classification of Proteins in a Diverse Group

Published on: August 16, 2017

7.4K
A Practical Guide to Phylogenetics for Nonexperts
12:00

A Practical Guide to Phylogenetics for Nonexperts

Published on: February 5, 2014

35.9K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Genomics

Background:

  • Alignment-free tools, often k-mer based, offer computational efficiency for high-throughput sequencing data analysis.
  • Traditional k-mer methods struggle with sequencing errors and polymorphisms, limiting sensitivity.
  • Spaced seeds improve mismatch tolerance but incur high computational and memory costs, hindering practical classification applications.

Purpose of the Study:

  • To develop a novel probabilistic data structure, the multi-index Bloom Filter (miBF), to overcome the limitations of existing spaced seed methods in bioinformatics.
  • To enable efficient storage and classification of multiple spaced seed sequences with low, static memory footprint.
  • To formalize false-positive rate minimization for miBFs in multi-target or multi-reference classification scenarios.

Main Methods:

  • Designed and implemented a multi-index Bloom Filter (miBF) as a probabilistic data structure.
  • Integrated miBF into BioBloom Tools for practical applications.
  • Evaluated miBF performance in read-binning for targeted assembly and taxonomic read assignment using benchmark analyses.

Main Results:

  • The miBF-based pipeline demonstrated superior sensitivity and specificity for read-binning compared to alignment-based methods, with faster execution times.
  • For taxonomic classification, miBF achieved higher sensitivity than conventional spaced seed approaches.
  • The miBF approach utilized half the memory and an order of magnitude less computational time compared to existing methods.

Conclusions:

  • The multi-index Bloom Filter (miBF) offers a memory-efficient and computationally fast solution for sequence classification in bioinformatics.
  • miBF significantly enhances the performance of alignment-free tools, particularly in read-binning and taxonomic assignment tasks.
  • This novel data structure addresses key challenges in spaced seed implementation, paving the way for more sensitive and scalable bioinformatics analyses.