Related Experiment Video
Updated: Jun 18, 2026

A Simple, Robust, and High Throughput Single Molecule Flow Stretching Assay Implementation for Studying Transport of Molecules Along DNA
Published on: October 1, 2017
Fast multiple alignment of ungapped DNA sequences using information theory and a relaxation method.
Thomas D Schneider1, David N Mastronarde
1National Cancer Institute, Frederick Cancer Research and Development Center, Laboratory of Mathematical Biology, P. O. Box B, Frederick, MD 21702-1201. toms@ncifcrf.gov.
A novel information theory method, Malign, accurately aligns challenging DNA binding sequences for OxyR and Fis proteins. This approach uses information content as a global quality criterion, enabling unambiguous selection of the best alignment.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Identifying DNA binding sites for regulatory proteins like OxyR and Fis can be challenging due to dispersed sequence conservation.
- Traditional alignment methods struggle with identifying functionally relevant sites in such cases.
Purpose of the Study:
- To develop and present an information theory-based multiple sequence alignment method (Malign) for improved identification of DNA binding sites.
- To establish a robust and unambiguous criterion for evaluating the quality of sequence alignments.
Main Methods:
- Utilized an information theory-based multiple alignment algorithm (Malign) to analyze DNA binding sequences of OxyR and Fis proteins.
- Employed information content as a global quality criterion for alignment, using look-up tables for computational efficiency.
- Applied a hill-climbing algorithm to navigate the vast combinatorial space of possible alignments, starting from random selections.
Main Results:
- The Malign method provides an unambiguous way to select the best alignment using absolute units (bits) without arbitrary constants.
- The algorithm's speed allows for multiple starting points and classification of solutions, indicating convergence by a single, well-populated high-information content class.
- Distinct solution classes for Fis protein suggest self-similar features within its DNA binding sites.
Conclusions:
- The Malign algorithm offers a powerful and objective approach for multiple sequence alignment, particularly for identifying difficult-to-discern DNA binding sites.
- The method's reliance on information content provides a reliable metric for assessing alignment quality and biological relevance.
- Analysis of Fis protein binding sites reveals potential self-similar structural or functional characteristics.
Related Concept Videos
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Sanger Sequencing
Evolutionary Relationships through Genome Comparisons
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...

