Related Experiment Video
Updated: Jun 19, 2026

09:14
Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
Published on: June 28, 2018
Characterizing the D2 statistic: word matches in biological sequences
Sylvain Forêt1, Susan R Wilson, Conrad J Burden
1Australian National University & James Cook University. sylvain.foret@anu.edu.au
Statistical Applications in Genetics and Molecular Biology
|November 4, 2009
Summary
The D2 statistic, used for sequence comparison, now accurately approximates distributions for approximate word matches. This enhances accuracy in biological sequence analysis and database searches.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Word matches are fundamental in sequence comparison algorithms like BLAST and BLAT.
- The D2 statistic quantifies word matches between sequences, crucial for similarity assessment.
- Existing methods for D2 statistic distribution approximation are being advanced.
Purpose of the Study:
- To extend D2 statistic characterization to approximate word matches.
- To compute the exact variance for uniform letter distributions and approximate it for others.
- To improve the precision of sequence comparison and biological data analysis.
Main Methods:
- Calculating the exact variance of the D2 statistic under uniform letter distribution.
- Developing accurate approximation methods for D2 statistic variance in non-uniform distributions.
- Applying the enhanced D2 statistic approximation to identify cis-regulatory modules.
Main Results:
- Exact variance computation for D2 statistic with uniform letter distribution.
- Accurate approximation methods for D2 statistic variance in various biological sequence contexts.
- High accuracy demonstrated in identifying cis-regulatory modules using the improved D2 statistic.
Conclusions:
- The enhanced D2 statistic approximation improves sequence comparison and database search precision.
- This method offers a more accurate way to identify biological sequences, including transcription factor binding sites.
- The findings facilitate more precise applications of the D2 statistic in bioinformatics research.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Wilcoxon Signed-Ranks Test for Matched Pairs
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
Wald-Wolfowitz Runs Test II
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
Wald-Wolfowitz Runs Test I
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
Sign Test for Matched Pairs
The sign test for matched pairs offers a robust method for comparing two paired samples, often for the effects of an intervention in one of them. This method is very useful in situations where the underlying distribution of the data is unknown. The test compares two related samples—often pre- and post-treatment measurements on the same subjects—to determine if there are significant differences in their median values.
To conduct the sign test, we first calculate the differences in value between...
To conduct the sign test, we first calculate the differences in value between...
