Related Experiment Video
Updated: May 22, 2026

12:00
A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
Alignment-free sequence comparison for biologically realistic sequences of moderate length
Conrad J Burden1, Junmei Jing, Susan R Wilson
1Australian National University.
Summary
The D(2) statistic and its variants D(2)* and D(2)c offer fast, alignment-free methods for assessing biological sequence similarity. These methods effectively identify cis-regulatory modules, with D(2) and D(2)c showing strong performance.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- The D(2) statistic is a computationally efficient, alignment-free method for measuring biological sequence similarity.
- Concerns exist regarding the D(2) statistic's suitability due to potential noise dominance from individual sequences.
- Mean-centered variants, D(2)* and D(2)c, were developed to address these concerns.
Purpose of the Study:
- To evaluate the suitability of the D(2) statistic for biological sequence similarity analysis.
- To assess the effectiveness of D(2)* and D(2)c in mitigating noise-related variability.
- To determine the utility of these statistics in identifying cis-regulatory modules.
Main Methods:
- Comparison of D(2), D(2)*, and D(2)c statistics.
- Examination of statistical variability and noise impact.
- Performance evaluation in classifying cis-regulatory modules using p-value estimation under a null hypothesis.
Main Results:
- All three statistics (D(2), D(2)*, D(2)c) are potentially useful for sequence similarity.
- Accurate p-values can be estimated for these statistics assuming independent and identically distributed sequence letters.
- D(2) and D(2)c demonstrated strong performance in classifying cis-regulatory modules, with D(2)* showing moderate effectiveness.
Conclusions:
- The D(2) statistic and its variants are valuable tools for biological sequence analysis.
- These alignment-free methods, particularly D(2) and D(2)c, show promise for identifying functional genomic elements like cis-regulatory modules.
- The study validates the utility of these statistics and their p-value estimation under specific hypotheses.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Multi-species Conserved Sequences
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved DNA...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...

