Accurate and highly interpretable prediction of gene expression from histone modifications
Fabrizio Frasca1,2, Matteo Matteucci3, Michele Leone3
1Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy. f.frasca18@imperial.ac.uk.
Background:
Histone Mark Modifications (HMs) are crucial actors in gene regulation, as they actively remodel chromatin to modulate transcriptional activity: aberrant combinatorial patterns of HMs have been connected with several diseases, including cancer. HMs are, however, reversible modifications: understanding their role in disease would allow the design of 'epigenetic drugs' for specific, non-invasive treatments. Standard statistical techniques were not entirely successful in extracting representative features from raw HM signals over gene locations. On the other hand, deep learning approaches allow for effective automatic feature extraction, but at the expense of model interpretation.
Results:
Here, we propose ShallowChrome, a novel computational pipeline to model transcriptional regulation via HMs in both an accurate and interpretable way. We attain state-of-the-art results on the binary classification of gene transcriptional states over 56 cell-types from the REMC database, largely outperforming recent deep learning approaches. We interpret our models by extracting insightful gene-specific regulative patterns, and we analyse them for the specific case of the PAX5 gene over three differentiated blood cell lines. Finally, we compare the patterns we obtained with the characteristic emission patterns of ChromHMM, and show that ShallowChrome is able to coherently rank groups of chromatin states w.r.t. their transcriptional activity.
Conclusions:
In this work we demonstrate that it is possible to model HM-modulated gene expression regulation in a highly accurate, yet interpretable way. Our feature extraction algorithm leverages on data downstream the identification of enriched regions to retrieve gene-wise, statistically significant and dynamically located features for each HM. These features are highly predictive of gene transcriptional state, and allow for accurate modeling by computationally efficient logistic regression models. These models allow a direct inspection and a rigorous interpretation, helping to formulate quantifiable hypotheses.
More Related Videos
08:12Global Level Quantification of Histone Post-Translational Modifications in a 3D Cell Culture Model of Hepatic Tissue
Published on: May 5, 2022
10:41An Integrated Platform for Genome-wide Mapping of Chromatin States Using High-throughput ChIP-sequencing in Tumor Tissues
Published on: April 5, 2018
Related Concept Videos
Histone Modification
Acetylation
The enzyme histone acetyltransferase adds acetyl group to the histones. Another enzyme, histone...
Spreading of Chromatin Modifications
Writers
The writer...
Chromatin Structure Regulates pre-mRNA Processing
The chromatin structure, especially...
Chromatin Position Affects Gene Expression
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
What is Gene Expression?
Chromatin Immunoprecipitation- ChIP
Types of ChIP
ChIP can be divided into two types - X-ChIP and N-ChIP. X-ChIP involves in vivo cross-linking of histones and regulatory proteins to DNA, fragmenting the DNA by sonication, and isolating the protein-DNA...
