Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Poisson Probability Distribution01:09

Poisson Probability Distribution

A Poisson probability distribution is a discrete probability distribution. It gives the probability of a number of events occurring in a fixed interval of time or space if these events happen at a known average rate and independently of the time since the last event. For example, a book editor might be interested in the number of words spelled incorrectly in a particular book. It might be that, on average, there are five words spelled incorrectly in 100 pages. The interval is 100 pages.
The...
Poisson's Ratio01:23

Poisson's Ratio

Poisson's ratio is a material property that indicates their stress response. It explains the connection between the elongation or compression a material undergoes in the direction of an applied force and the contraction or expansion it experiences perpendicular to that force. When a slender bar is loaded axially, it stretches in the direction of the force and contracts laterally. Poisson's ratio is the negative ratio of this lateral contraction to the axial elongation. The negative sign ensures...
Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
Poisson's And Laplace's Equation01:25

Poisson's And Laplace's Equation

The electric potential of the system can be calculated by relating it to the electric charge densities that give rise to the electric potential. The differential form of Gauss's law expresses the electric field's divergence in terms of the electric charge density.
Real Time RT-PCR02:57

Real Time RT-PCR

Real-time reverse transcription-polymerase chain reaction, or Real-time RT-PCR, is an analytical tool used to determine the expression level of target genes. The method involves converting mRNA to complementary DNA with the help of an enzyme known as reverse transcriptase, followed by the PCR amplification of the cDNA. These two processes can be performed simultaneously in a single tube or separately as a two-step reaction.
The real-time quantification of the number of amplified products is...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Lethal conflict after group fission in wild chimpanzees.

Science (New York, N.Y.)·2026
Same author

Metagenomic Hi‑C Protocols for Viral Genome Binning, Taxonomic Annotation, and Interaction Network Visualization.

Current protocols·2026
Same author

Benchmarking alignment strategies for Hi-C reads in metagenomic Hi-C data.

Genome biology·2026
Same author

A Beginner's Guide to Using DeepVirFinder for Viral Sequence Identification From Metagenomic Datasets.

Current protocols·2026
Same author

Correction: Quantifying microbial interactions based on compositional data using an iterative approach for solving generalized Lotka-Volterra equations.

PLoS computational biology·2026
Same author

ViTrace detects viral signatures in tumor transcriptomes using a hybrid language model.

Communications biology·2025

Related Experiment Video

Updated: May 21, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

Normal and compound poisson approximations for pattern occurrences in NGS reads.

Zhiyuan Zhai1, Gesine Reinert, Kai Song

  • 1School of Mathematics, Shandong University, Jinan, Shandong, China.

Journal of Computational Biology : a Journal of Computational Molecular Cell Biology
|June 16, 2012
PubMed
Summary

This study introduces a novel word pattern analysis for next-generation sequencing (NGS) data, offering accurate approximations for motif discovery even with unknown genomes. The compound Poisson approximation generally outperforms normal approximation for analyzing sequence reads.

More Related Videos

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

Pattern-based Search of Epigenomic Data Using GeNemo
06:38

Pattern-based Search of Epigenomic Data Using GeNemo

Published on: October 8, 2017

Related Experiment Videos

Last Updated: May 21, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

Pattern-based Search of Epigenomic Data Using GeNemo
06:38

Pattern-based Search of Epigenomic Data Using GeNemo

Published on: October 8, 2017

Area of Science:

  • Computational Biology
  • Bioinformatics
  • Genomics

Background:

  • Next-generation sequencing (NGS) generates vast amounts of sequence reads, often challenging to analyze due to unknown genomes or unmappable reads.
  • Traditional NGS data analysis relies on mapping reads to a reference genome, which is not always feasible.
  • Word pattern counting is a powerful tool in molecular sequence analysis, but its application to NGS reads requires specific probabilistic modeling.

Purpose of the Study:

  • To develop and evaluate novel computational methods for analyzing next-generation sequencing (NGS) data using word patterns.
  • To provide accurate probabilistic approximations for motif discovery in NGS reads, addressing limitations of traditional mapping-based approaches.
  • To assess the performance of normal and compound Poisson approximations for quantifying word pattern occurrences in NGS data.

Main Methods:

  • Development of probabilistic models for background genome sequences and the NGS read sampling process.
  • Derivation of normal and compound Poisson approximations for the number of occurrences of word patterns in NGS reads.
  • Evaluation of approximation accuracy under various conditions and comparison with traditional mapping methods.

Main Results:

  • The study successfully builds probabilistic models and provides accurate normal and compound Poisson approximations for word pattern occurrences in NGS reads.
  • Compound Poisson approximation demonstrates superior performance over normal approximation in most realistic scenarios for analyzing sequence reads.
  • The developed methods and algorithms were applied to analyze ChIP-Seq data for the transcription factor GABP.

Conclusions:

  • Word pattern analysis offers a viable alternative for NGS data analysis, particularly when reference genomes are unavailable or reads are unmappable.
  • The compound Poisson approximation provides a robust statistical framework for evaluating the significance of sequence patterns in NGS data.
  • The developed computational tools and theoretical framework enhance the capabilities for motif discovery and biological interpretation of NGS experiments.