Related Experiment Video
Updated: Jul 25, 2026

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
Published on: April 1, 2016
Probabilistic and statistical properties of words: an overview
G Reinert1, S Schbath, M S Waterman
1King's College and Statistical Laboratory, Cambridge, UK.
This study analyzes statistical properties of words in biological sequences using Markov chains. It develops methods for accurate pattern analysis and confidence intervals, improving sequence data interpretation.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Genomics
Background:
- Analyzing biological sequences involves understanding word occurrences and their statistical properties.
- Accurate statistical models are crucial for interpreting complex biological sequence data.
Purpose of the Study:
- To provide an overview of statistical and probabilistic properties of words in biological sequence analysis.
- To develop methods for analyzing word occurrences, including exact distributions and various approximations.
- To address the challenges of dependence structures between word occurrences and sequence recoverability.
Main Methods:
- Modeling biological sequences as stationary ergodic Markov chains.
- Deriving exact distributions and approximations (normal, Poisson, compound Poisson) for word counts.
- Utilizing moment generating functions, martingales, Stein's method, and the Chen-Stein method.
- Developing a test for determining Markov chain order and accounting for estimation errors.
Main Results:
- Distinguished counts of occurrence, clumps, and renewals with derived distributions and approximations.
- Established convergence results considering estimation errors in Markovian transition probabilities.
- Provided methods for analyzing multiple pattern occurrences and sequence recoverability from SBH chip data.
- Disentangled complex dependence structures between word occurrences due to self-overlap and inter-word overlap.
Conclusions:
- The developed statistical methods offer accurate and conservative confidence intervals for hypothesis testing in biological sequence analysis.
- The study enhances the understanding of word patterns in biological sequences, aiding in data interpretation.
- The findings are applicable to various bioinformatics problems, including sequence assembly and analysis of high-throughput sequencing data.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Related Concept Videos
Probability in Statistics
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
Random Variables
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Probability Distributions
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson probability...
Probability Histograms
Poisson Probability Distribution
The...
Biostatistics: Overview
Discrete variables are...