Related Experiment Video
Updated: Jun 25, 2026

08:04
DNA Sequence Recognition by DNA Primase Using High-Throughput Primase Profiling
Published on: October 8, 2019
A reexamination of information theory-based methods for DNA-binding site identification.
Ivan Erill1, Michael C O'Neill
1Department of Biological Sciences, University of Maryland-Baltimore County, Baltimore, MD, USA. erill@umbc.edu
BMC Bioinformatics
|February 13, 2009
Summary
Information theory methods for finding transcription factor binding sites often overestimate efficiency on real genomes. Simple methods work best, as complex ones fail, revealing new insights into binding site evolution.
Area of Science:
- Bioinformatics
- Genomics
- Computational Biology
Background:
- Transcription factor binding site (TFBS) identification remains a challenge in bioinformatics.
- Information theory (IT) based search methods are standard but often validated on artificial data.
- This study assesses IT-based TFBS search methods using real bacterial genomic data.
Purpose of the Study:
- To evaluate the performance of IT-based TFBS search methods on real genomic data.
- To identify limitations of current methods and propose a revised understanding of TFBS information content.
- To reassess the validity of IT assumptions in the context of biological sequences.
Main Methods:
- Benchmarking IT-based TFBS search methods using diverse bacterial genome datasets.
- Analyzing the impact of genomic features like sequence skew and curvature on search efficiency.
- Comparing performance of simple vs. heuristically corrected IT search algorithms.
Main Results:
- Conventional benchmarking on artificial data overestimates TFBS search efficiency.
- Sequence information alone is insufficient; genomic features like curvature are crucial.
- Methods integrating genomic skew information, like Relative Entropy, are ineffective on real genomes.
- Binding sites evolve towards genomic skew, maintaining information via conservation.
- Identified misconceptions regarding IT application to TFBS, proposing a revised paradigm.
Conclusions:
- Unassuming IT-based TFBS search methods outperform complex alternatives on real data.
- Heuristic corrections to IT methods often fail with biological sequences.
- TFBS information content is a composite measure of search and binding affinity requirements.
- This understanding has significant implications for TFBS evolution studies.
Related Concept Videos
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservative Site-specific Recombination and Phase Variation
Because the DNA segments are cut and reorganized in a direction-specific manner, site-specific recombination has emerged as an efficient genetic engineering technique. Flippase and Cyclization recombinases or Flp and Cre, respectively, are two members of the tyrosine recombinase family derived from bacteriophages, that are used to mediate site-specific DNA insertions, deletions, and targeted expression of proteins in mammalian cell lines.
The recognition sites for Cre recombinase called LoxP...
The recognition sites for Cre recombinase called LoxP...
Ligand Binding Sites
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...

