Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Unusual Results01:16

Unusual Results

3.7K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ  from the mean, μ  is considered unusual.
Maximum unusual value =...
3.7K
Probability in Statistics01:14

Probability in Statistics

21.9K
Probability is the likelihood of an event occurring. The term event is defined as a collection of results of a procedure. An event is a simple event when an outcome cannot be divided into simpler parts.
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
21.9K
Probability Laws01:49

Probability Laws

43.8K
Overview
43.8K
Wald-Wolfowitz Runs Test I01:17

Wald-Wolfowitz Runs Test I

922
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
922
Random Variables01:09

Random Variables

17.2K
A random variable is a single numerical value that indicates the outcome of a procedure. The concept of random variables is fundamental to the probability theory and was introduced by a Russian mathematician, Pafnuty Chebyshev, in the mid-nineteenth century.
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
17.2K
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

7.0K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
7.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Introgression dynamics of sex-linked chromosomal inversions shape the Malawi cichlid radiation.

Science (New York, N.Y.)·2025
Same author

Ensembl 2009.

Nucleic acids research·2008
Same author

Ensembl 2008.

Nucleic acids research·2007
Same author

Ensembl 2007.

Nucleic acids research·2006
Same author

The DNA sequence and biological annotation of human chromosome 1.

Nature·2006
Same author

Ensembl 2006.

Nucleic acids research·2005

Related Experiment Video

Updated: Jan 9, 2026

Exploring Sequence Space to Identify Binding Sites for Regulatory RNA-Binding Proteins
11:34

Exploring Sequence Space to Identify Binding Sites for Regulatory RNA-Binding Proteins

Published on: August 9, 2019

7.1K

Method for calculation of probability of matching a bounded regular expression in a random data string

R F Sewell1, R Durbin

  • 1Sanger Centre, Hinxton, Cambridge, UK.

Journal of Computational Biology : a Journal of Computational Molecular Cell Biology
|January 1, 1995
PubMed
Summary

This study introduces a method to calculate the probability of a regular expression matching a specific point in random data. The technique accurately bounds these probabilities, even for complex patterns like ProSite sequences.

More Related Videos

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.5K
Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
09:23

Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans

Published on: August 16, 2017

8.5K

Related Experiment Videos

Last Updated: Jan 9, 2026

Exploring Sequence Space to Identify Binding Sites for Regulatory RNA-Binding Proteins
11:34

Exploring Sequence Space to Identify Binding Sites for Regulatory RNA-Binding Proteins

Published on: August 9, 2019

7.1K
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
10:36

Rare Event Detection Using Error-corrected DNA and RNA Sequencing

Published on: August 3, 2018

12.5K
Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans
09:23

Quantification of Information Encoded by Gene Expression Levels During Lifespan Modulation Under Broad-range Dietary Restriction in C. elegans

Published on: August 16, 2017

8.5K

Area of Science:

  • Computational biology
  • Bioinformatics
  • Algorithm analysis

Background:

  • Probabilistic matching of patterns in biological sequences is crucial for data analysis.
  • Accurate estimation of match probabilities aids in sequence annotation and database searching.
  • Existing methods may struggle with the complexity of biological pattern matching.

Purpose of the Study:

  • To develop a method for precisely determining the probability of a regular expression match at a specific start point within a random data string.
  • To establish strict bounds for these probabilities, ensuring reliability in computational analyses.
  • To assess the method's applicability to complex biological pattern databases.

Main Methods:

  • The core method involves calculating the probability of a regular expression match within defined boundaries.
  • It addresses the computational complexity, which can be exponential relative to the number of optional characters in the expression.
  • The approach was practically applied to evaluate probabilities for all ProSite patterns.

Main Results:

  • The method successfully determined strict bounds for the probability of regular expression matches.
  • Despite theoretical exponential complexity, the practical application to ProSite patterns was achieved without significant difficulty.
  • This demonstrates the method's efficacy for real-world biological sequence analysis.

Conclusions:

  • The presented method offers a robust way to quantify the certainty of pattern matches in random data.
  • It is particularly valuable for analyzing complex biological patterns, such as those found in the ProSite database.
  • This work advances the accuracy and reliability of computational tools used in bioinformatics and sequence analysis.