Epigenetic priors for identifying active transcription factor binding sites
Gabriel Cuellar-Partida1, Fabian A Buske, Robert C McLeay
1Institute for Molecular Bioscience, The University of Queensland, Brisbane QLD 4072, Australia.
Bioinformatics (Oxford, England)
|November 11, 2011
Summary
This study introduces a new probabilistic method to identify active transcription factor binding sites (TFBSs) by integrating epigenetic data with DNA sequence motifs. The method outperforms existing filters and offers a simpler, general approach for TFBS prediction.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Understanding genome-wide transcription factor binding is crucial for transcriptional regulation.
- Epigenetic data, including histone modifications and DNase I sensitivity, can enhance in silico prediction of transcription factor binding sites (TFBSs).
Purpose of the Study:
- To develop and validate a probabilistic method for improving the identification of active TFBSs by integrating epigenetic data with DNA sequence motif models.
- To compare the performance of the proposed method against existing approaches.
Main Methods:
- A probabilistic method was developed to combine multiple epigenetic data tracks (H3K4me1, H3K4me3, H3K9ac, H3K27ac, DNase I sensitivity) with DNA sequence motif models.
- Each data type was converted into a position-specific probabilistic prior, which was then combined with a traditional motif model to compute a log-posterior odds score.
- The FIMO software, part of the MEME Suite, was updated to support this log-posterior odds scoring.
Main Results:
- The log-posterior odds score consistently outperformed a simple binary filter using the same epigenetic data.
- The proposed method demonstrated competitive performance compared to the more complex CENTIPEDE method.
- The method effectively identifies active transcription factor binding sites using DNA sequence and epigenetic evidence.
Conclusions:
- The developed probabilistic method provides a robust and general approach for identifying functional TFBSs.
- The simplicity of the log-posterior odds scoring makes it an appealing tool for researchers in transcriptional regulation.
- Integration of epigenetic data significantly improves the accuracy of in silico TFBS prediction.
Related Concept Videos
Transcription Factors
Tissue-specific transcription factors contribute to diverse cellular functions in mammals. For example, the gene for beta globin, a major component of hemoglobin, is present in all cells of the body. However, it is only expressed in red blood cells because the transcription factors that can bind to the promoter sequences of the beta globin gene are only expressed in these cells. Tissue-specific transcription factors also ensure that mutations in these factors may impair only the function of...
Transcription Factors
Tissue-specific transcription factors contribute to diverse cellular functions in mammals. For example, the gene for beta globin, a major component of hemoglobin, is present in all cells of the body. However, it is only expressed in red blood cells because the transcription factors that can bind to the promoter sequences of the beta globin gene are only expressed in these cells. Tissue-specific transcription factors also ensure that mutations in these factors may impair only the function of...
Chromatin Immunoprecipitation- ChIP
Chromatin immunoprecipitation, or ChIP, is an antibody-based technique used to identify sites on DNA that bind to transcription factors of interest or histone proteins. It also helps determine the type of histone modifications such as acetylation, phosphorylation, or methylation.
Types of ChIP
ChIP can be divided into two types - X-ChIP and N-ChIP. X-ChIP involves in vivo cross-linking of histones and regulatory proteins to DNA, fragmenting the DNA by sonication, and isolating the protein-DNA...
Types of ChIP
ChIP can be divided into two types - X-ChIP and N-ChIP. X-ChIP involves in vivo cross-linking of histones and regulatory proteins to DNA, fragmenting the DNA by sonication, and isolating the protein-DNA...
RNA Polymerase II Accessory Proteins
Proteins that regulate transcription can do so either via direct contact with RNA Polymerase or through indirect interactions facilitated by adaptors, mediators, histone-modifying proteins, and nucleosome remodelers. Direct interactions to activate transcription is seen in bacteria as well as in some eukaryotic genes. In these cases, upstream activation sequences are adjacent to the promoters, and the activator proteins interact directly with the transcriptional machinery. For example, in...
Co-activators and Co-repressors
Gene transcription is regulated by the synergistic action of several proteins that form a complex at a gene regulatory site. This is observed in eukaryotes, where the regulation of gene expression is a complex process. Regulatory proteins in eukaryotes can broadly be classified into two types – regulators that bind directly to specific DNA sequences and co-regulators that associate with regulatory proteins but cannot directly bind to the DNA. These co-regulators are further divided into...
Eukaryotic Transcription Activators
Transcription activators are proteins that promote the transcription of genes from DNA to RNA. In most cases, these proteins contain two separate domains ‒ a domain that binds to DNA and a domain for activating transcription; however, in some cases, a single domain is responsible for both binding and activation of transcription, as seen in the glucocorticoid receptor and MyoD.
The binding domains are capable of recognizing and interacting with regulatory sequences on the DNA. These domains are...
The binding domains are capable of recognizing and interacting with regulatory sequences on the DNA. These domains are...


