Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Unusual Results01:16

Unusual Results

Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ  from the mean, μ  is considered unusual.
Maximum unusual value = μ + 2σ
Minimum unusual value...
Wald-Wolfowitz Runs Test II01:17

Wald-Wolfowitz Runs Test II

The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
Finding Critical Values for Chi-Square01:18

Finding Critical Values for Chi-Square

Consider a curve representing sample data drawn randomly from a normally distributed population. One must construct confidence intervals to estimate or to test a claim regarding the population standard deviation. For example, a 95% confidence interval covers 95% of the area under the curve, and the remaining 5% is equally distributed on either side of the curve. To achieve such confidence intervals, one must determine the critical values. The critical values are simply the values separating the...
Random Error01:04

Random Error

Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Numerical estimation of limiting large-deviation rate functions.

Physical review. E·2026
Same author

Rare events of host switching for diseases using a susceptible-infected-recovered model with mutations.

Physical review. E·2026
Same author

Distribution of the Number of Paths in Two-Dimensional Directed Percolation.

Entropy (Basel, Switzerland)·2025
Same author

Diffusion with stochastic resetting on a lattice.

Physical review. E·2025
Same author

Nonuniversality for crossword puzzle percolation.

Physical review. E·2025
Same author

Resetting by rescaling: Exact results for a diffusing particle in one dimension.

Physical review. E·2024

Related Experiment Video

Updated: Jul 13, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

Local sequence alignments statistics: deviations from Gumbel statistics in the rare-event tail.

Stefan Wolfsheimer1, Bernd Burghardt, Alexander K Hartmann

  • 1Institut für Theoretische Physik, Universität Göttingen, 37077, Göttingen, Friedrich-Hund-Platz 1, Germany. wolfsh@theorie.physik.uni-oldenburg.de

Algorithms for Molecular Biology : AMB
|July 13, 2007
PubMed
Summary

The statistical distribution of gapped local sequence alignments deviates from the Gumbel distribution in biologically relevant rare events. A modified Gumbel distribution with a Gaussian factor improves accuracy for protein alignment significance estimations.

More Related Videos

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

Quantitation and Analysis of the Formation of HO-Endonuclease Stimulated Chromosomal Translocations by Single-Strand Annealing in Saccharomyces cerevisiae
09:40

Quantitation and Analysis of the Formation of HO-Endonuclease Stimulated Chromosomal Translocations by Single-Strand Annealing in Saccharomyces cerevisiae

Published on: September 23, 2011

Related Experiment Videos

Last Updated: Jul 13, 2026

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

Quantitation and Analysis of the Formation of HO-Endonuclease Stimulated Chromosomal Translocations by Single-Strand Annealing in Saccharomyces cerevisiae
09:40

Quantitation and Analysis of the Formation of HO-Endonuclease Stimulated Chromosomal Translocations by Single-Strand Annealing in Saccharomyces cerevisiae

Published on: September 23, 2011

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Statistical Modeling

Background:

  • Optimal scores for ungapped local sequence alignments follow a Gumbel distribution.
  • The distribution for gapped alignments is less understood, particularly in biologically relevant rare-event regions.

Purpose of the Study:

  • To develop a numerical method for determining the rare-event tail of gapped local alignment distributions.
  • To investigate the accuracy of the Gumbel distribution for gapped alignments and propose corrections.

Main Methods:

  • Utilized Metropolis Coupled Markov Chain Monte Carlo (MCMC) with biased probability distributions.
  • Generated sequences under parametrized distributions to simulate biological scenarios.
  • Analyzed distributions for protein alignments using various substitution matrices (BLOSUM62, PAM250) and affine gap costs.

Main Results:

  • The Gumbel distribution is insufficient for describing the rare-event tail of gapped alignments.
  • A modified Gumbel distribution, incorporating a Gaussian factor, accurately models these distributions for sequences up to length 400.
  • Significance estimations differ considerably when using these refined distributions compared to standard BLAST parameters.

Conclusions:

  • Gapped and ungapped local alignment statistics diverge from the Gumbel distribution in rare-event tails.
  • A Gaussian correction to the Gumbel distribution is proposed and its scaling behavior analyzed for common protein database search parameters.
  • The study also addresses the distribution of sum statistics for the k-best alignments.