Integration of genome and chromatin structure with gene expression profiles to predict c-MYC recognition site binding

Yili Chen1, Thomas W Blackwell, Ji Chen

  • 1Bioinformatics Program, University of Michigan Medical School, Ann Arbor, Michigan, United States of America.

Insights

A new computational model accurately predicts MYC gene targets by integrating genomic data, improving cancer research. This approach identifies novel MYC targets, advancing our understanding of cellular regulation in cancer.

Area of Science:

  • Genomics
  • Molecular Biology
  • Bioinformatics

Background:

  • MYC genes are crucial regulators of cellular functions and are frequently altered in human cancers.
  • Experimental identification of MYC gene targets is inconsistent, with many predicted sites lacking experimental validation.
  • Accurate identification of MYC targets is essential for understanding cancer development and for therapeutic strategies.

Purpose of the Study:

  • To develop a computational model for accurate prediction of in vivo MYC gene binding and regulation.
  • To integrate diverse data sources, including genomic sequence, chromatin accessibility, DNA methylation, and gene expression, for improved target prediction.
  • To identify novel MYC target genes and validate predictions against existing literature.

Main Methods:

  • Utilized a Bayesian network classifier incorporating genomic sequence, chromatin acetylation, and DNA methylation predictions.
  • Integrated predicted MYC binding probabilities with gene expression data from multiple microarray datasets.
  • Employed Gene Ontology annotations to refine predictions of MYC target genes.

Main Results:

  • Predicted 460 likely c-MYC target genes in the human genome.
  • Validated predictions against known MYC targets, identifying 67 bound and regulated, 68 bound, and 80 regulated genes.
  • Demonstrated the model's applicability to other transcription factors sensitive to DNA methylation.

Conclusions:

  • The integrated computational approach significantly improves the accuracy of identifying MYC genomic targets.
  • The study successfully identified numerous known MYC targets and proposed a substantial list of novel candidates.
  • This methodology provides a robust framework for discovering transcription factor targets in complex biological systems.

Related Concept Videos

Chromatin Immunoprecipitation- ChIP02:36

Chromatin Immunoprecipitation- ChIP

Chromatin immunoprecipitation, or ChIP, is an antibody-based technique used to identify sites on DNA that bind to transcription factors of interest or histone proteins. It also helps determine the type of histone modifications such as acetylation, phosphorylation, or methylation.
Types of ChIP
ChIP can be divided into two types - X-ChIP and N-ChIP. X-ChIP involves in vivo cross-linking of histones and regulatory proteins to DNA, fragmenting the DNA by sonication, and isolating the protein-DNA...
Chromatin Position Affects Gene Expression02:35

Chromatin Position Affects Gene Expression

Chromatin is the massive complex of DNA and proteins packaged inside the nucleus. The complexity of chromatin folding and how it is packaged inside the nucleus greatly influences  access to genetic information. Generally, the nucleus' periphery is considered transcriptionally repressive, while the cell's interior is considered a transcriptionally active area. 
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the timing and level of...
Induced Pluripotent Stem Cells01:06

Induced Pluripotent Stem Cells

Stem cells are undifferentiated cells that divide and produce different cell types. Ordinarily, cells that have differentiated into a specific cell type are terminally differentiated; however, scientists have found a way to reprogram these mature cells so that they dedifferentiate and return to an unspecialized, proliferative state. These cells are pluripotent like embryonic stem cells—able to produce all cell types—and are called induced pluripotent stem cells (iPSCs).
Somatic cells are...
Master Transcription Regulators02:23

Master Transcription Regulators

Master transcription regulators are regulatory proteins that are predominantly responsible for regulating the expression of multiple genes. Often these genes work in concert to drive a  complex process. Activation of a master transcription regulator can lead to a cascade of transcriptional activation necessary for that outcome. These regulators can directly bind to the regulatory sequences of the various genes involved, or they can indirectly regulate transcription by binding to regulatory...
Chromatin Structure Regulates pre-mRNA Processing02:41

Chromatin Structure Regulates pre-mRNA Processing

In eukaryotic cells, nascent mRNA transcripts need to undergo many post-transcriptional modifications to reach the cell cytoplasm and translate into functional proteins. For a long time, transcription and pre-mRNA processing were considered two independent events that occur sequentially in the cell. However, it has now been well established that transcription and pre-mRNA processing are two simultaneous processes that are precisely regulated inside the cell.
The chromatin structure, especially...
Combinatorial Gene Control02:33

Combinatorial Gene Control

Combinatorial gene control is the synergistic action of several transcriptional factors to regulate the expression of a single gene. The absence of one or more of these factors may lead to a significant difference in the level of gene expression or repression.
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...