Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genome Annotation and Assembly03:36

Genome Annotation and Assembly

The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Protein Complex Assembly02:41

Protein Complex Assembly

Proteins can form homomeric complexes with another unit of the same protein or heteromeric complexes with different types.  Most protein complexes self-assemble spontaneously via ordered pathways, while some proteins need assembly factors that guide their proper assembly. Despite the crowded intracellular environment, proteins usually interact with their correct partners and form functional complexes.
Many viruses self-assemble into a fully functional unit using the infected host cell to...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

AIEdit: Alignment-free genome assembly polisher trained on spaced seed match patterns.

PLoS computational biology·2026
Same author

ntStat: k-mer characterization using occurrence statistics in raw sequencing data.

PLoS computational biology·2026
Same author

A comprehensive tandem repeat catalog of the human genome.

Nature communications·2026
Same author

ntSynt: multi-genome synteny detection using minimizer graph mappings.

BMC biology·2025
Same author

ntRoot: computational inference of human ancestry at scale from genomic data.

Bioinformatics advances·2025
Same author

Antimicrobial peptides with high bioactivity against MDR isolates: Addressing public health concerns.

Microbial pathogenesis·2025

Related Experiment Video

Updated: Jul 15, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
08:03

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations

Published on: December 7, 2021

Genome misassembly detection using Stash: A data structure based on stochastic tile hashing.

Armaghan Sarvar1,2, Lauren Coombe2, Inanc Birol1,2

  • 1University of British Columbia, Vancouver, Canada.

Plos One
|July 13, 2026
PubMed
Summary

Stash, a new hash-based data structure, efficiently detects and corrects genome misassemblies in large sequencing data. This method improves the accuracy of de novo genome assembly, crucial for genomics research.

More Related Videos

Self-assembly of Complex Two-dimensional Shapes from Single-stranded DNA Tiles
10:23

Self-assembly of Complex Two-dimensional Shapes from Single-stranded DNA Tiles

Published on: May 8, 2015

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
09:14

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens

Published on: June 28, 2018

Related Experiment Videos

Last Updated: Jul 15, 2026

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
08:03

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations

Published on: December 7, 2021

Self-assembly of Complex Two-dimensional Shapes from Single-stranded DNA Tiles
10:23

Self-assembly of Complex Two-dimensional Shapes from Single-stranded DNA Tiles

Published on: May 8, 2015

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
09:14

Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens

Published on: June 28, 2018

Area of Science:

  • Bioinformatics
  • Genomics
  • Computational Biology

Background:

  • High-throughput sequencing generates large datasets, posing memory and computational challenges.
  • De novo genome assembly is fundamental to genomics but can suffer from misassemblies due to read errors and algorithmic limitations.
  • Accurate genome representation is vital for downstream analyses, necessitating improved assembly methods.

Purpose of the Study:

  • To introduce Stash, a novel hash-based data structure for efficient storage and querying of large sequencing data.
  • To leverage Stash for detecting and correcting misassemblies in de novo genome assemblies.
  • To evaluate Stash's performance in improving human genome assembly quality.

Main Methods:

  • Stash utilizes sliding windows and spaced seed patterns to extract and hash k-mers from input sequences.
  • Hash values and sequence IDs are stored in Stash to enable querying of read coverage across genomic regions.
  • The method was applied to detect misassemblies in human genome assemblies (Flye and Shasta) using Pacbio HiFi reads.

Main Results:

  • Stash effectively detected misassemblies in human genome assemblies.
  • Scaffolding with Stash reduced misassemblies by 7.6% in Flye assemblies and 3.4% in Shasta assemblies.
  • The process was computationally efficient, requiring 8 GB of memory and 310 minutes.

Conclusions:

  • Stash is a viable and efficient data structure for handling large sequencing data in bioinformatics.
  • Stash-based misassembly detection and correction significantly improves the quality of de novo genome assemblies.
  • Stash offers a competitive alternative to existing long-read misassembly correction methods, potentially yielding superior genomic data.