Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

14.4K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.4K
Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

11.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
11.5K
DNA Isolation01:24

DNA Isolation

35.3K
DNA isolation protocols can be fast and straightforward or complex and time-consuming depending on the type and quality of DNA required for further processing. For example, plasmid DNA extraction is a bit more complicated than genomic DNA extraction because of the need for an appropriate lysis method to separate plasmid DNA from gDNA during isolation. However, for specific applications, such as long-range DNA sequencing that require a good yield of high- quality DNA samples, we need to follow...
35.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Ill Fate of Rectal Mucinous Adenocarcinoma: A Defect in Immunosurveillance or a Mucin Coating Effect?-The IMMUNOREACT 20 Study.

Cancers·2026
Same author

N2SIMBA: from Network topology to SIMulation of interactions and BActerial abundance, using microbial consumer resource model.

Frontiers in bioinformatics·2026
Same author

Environmental Personal Exposure Clusters to Investigate Multiple Sclerosis and Amyotrophic Lateral Sclerosis Progression.

Studies in health technology and informatics·2026
Same author

IMMUNOREACT 4: Peritumoral Microenvironment Associated with Anastomotic Leaks After Surgery for Rectal Cancer.

Cancers·2026
Same author

The association of environmental exposure with multiple sclerosis severity score: A study based on sequential data modeling.

International journal of medical informatics·2026
Same author

MOV&RSim: computational modelling of cancer-specific variants and sequencing reads characteristics for realistic tumoral sample simulation.

BMC bioinformatics·2025

Related Experiment Video

Updated: Apr 26, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
09:34

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease

Published on: April 4, 2018

36.1K

Compression and fast retrieval of SNP data.

Francesco Sambo1, Barbara Di Camillo1, Gianna Toffolo1

  • 1Department of Information Engineering, University of Padova, via Gradenigo 6/a, 35131 Padova, Italy.

Bioinformatics (Oxford, England)
|July 28, 2014
PubMed
Summary

A new algorithm efficiently compresses and retrieves single nucleotide polymorphism (SNP) data for large genome-wide association studies. This method offers competitive compression rates and faster data loading times compared to existing tools.

More Related Videos

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

9.2K
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

16.5K

Related Experiment Videos

Last Updated: Apr 26, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
09:34

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease

Published on: April 4, 2018

36.1K
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

9.2K
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

16.5K

Area of Science:

  • Genetics
  • Bioinformatics
  • Computational Biology

Background:

  • Genome-wide association studies (GWAS) are expanding in size, necessitating efficient data handling.
  • Increasing focus on rare genetic variants and epistasis drives demand for large-scale genetic datasets.
  • Current methods struggle with the storage and retrieval demands of massive single nucleotide polymorphism (SNP) data.

Purpose of the Study:

  • To develop a novel algorithm and file format for compressing and rapidly retrieving SNP data.
  • To address the need for efficient data management in large-scale genetic association studies.
  • To improve the performance of SNP data compression and loading.

Main Methods:

  • A novel compression algorithm utilizing linkage disequilibrium (LD) blocks and reference SNP information.
  • Compression of LD blocks based on differences from a reference SNP.
  • Exploitation of call rate and minor allele frequency for reference SNP compression.

Main Results:

  • The developed algorithm achieves competitive compression rates for SNP data.
  • The algorithm significantly outperforms existing tools in terms of compressed data loading time.
  • Demonstrated effectiveness on two real-world SNP datasets.

Conclusions:

  • The novel algorithm provides an efficient solution for SNP data compression and retrieval.
  • This method is well-suited for the demands of large-scale genome-wide association studies.
  • The C++ library implementation is freely available, promoting its adoption.