Learning Useful Representations of DNA Sequences From ChIP-Seq Datasets for Exploring Transcription Factor Binding
Abstract:
Deep learning has been successfully applied to surprisingly different domains. Researchers and practitioners are employing trained deep learning models to enrich our knowledge. Transcription factors (TFs)are essential for regulating gene expression in all organisms by binding to specific DNA sequences. Here, we designed a deep learning model named SemanticCS (Semantic ChIP-seq)to predict TF binding specificities. We trained our learning model on an ensemble of ChIP-seq datasets (Multi-TF-cell)to learn useful intermediate features across multiple TFs and cells. To interpret these feature vectors, visualization analysis was used. Our results indicate that these learned representations can be used to train shallow machines for other tasks. Using diverse experimental data and evaluation metrics, we show that SemanticCS outperforms other popular methods. In addition, from experimental data, SemanticCS can help to identify the substitutions that cause regulatory abnormalities and to evaluate the effect of substitutions on the binding affinity for the RXR transcription factor. The online server for SemanticCS is freely available at http://qianglab.scst.suda.edu.cn/semanticCS/.
More Related Videos
09:52Generation of High Quality Chromatin Immunoprecipitation DNA Template for High-throughput Sequencing ChIP-seq
Published on: April 19, 2013
06:38High Sensitivity Measurement of Transcription Factor-DNA Binding Affinities by Competitive Titration Using Fluorescence Microscopy
Published on: February 7, 2019
Related Concept Videos
Cooperative Binding of Transcription Regulators
Cooperative Binding of Transcription Regulators
Chromatin Immunoprecipitation- ChIP
Types of ChIP
ChIP can be divided into two types - X-ChIP and N-ChIP. X-ChIP involves in vivo cross-linking of histones and regulatory proteins to DNA, fragmenting the DNA by sonication, and isolating the protein-DNA...
Transcription Factors
Cis-regulatory Sequences
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
