Epigene functional diversity: isoform usage, disordered domain content, and variable binding partners
Leroy Bondhus1,2,3, Aileen A Nava1,2,3, Isabelle S Liu1,2,3
1Department of Human Genetics, David Geffen School of Medicine, UCLA, 615 Charles E. Young Drive South, Los Angeles, CA, 90095, USA.
Background:
Epigenes are defined as proteins that perform post-translational modification of histones or DNA, reading of post-translational modifications, form complexes with epigenetic factors or changing the general structure of chromatin. This specialized group of proteins is responsible for controlling the organization of genomic DNA in a cell-type specific fashion, controlling normal development in a spatial and temporal fashion. Moreover, mutations in epigenes have been implicated as causal in germline pediatric disorders and as driver mutations in cancer. Despite their importance to human disease, to date, there has not been a systematic analysis of the sources of functional diversity for epigenes at large. Epigenes' unique functions that require the assembly of pools within the nucleus suggest that their structure and amino acid composition would have been enriched for features that enable efficient assembly of chromatin and DNA for transcription, splicing, and post-translational modifications.
Results:
In this study, we assess the functional diversity stemming from gene structure, isoforms, protein domains, and multiprotein complex formation that drive the functions of established epigenes. We found that there are specific structural features that enable epigenes to perform their variable roles depending on the cellular and environmental context. First, epigenes are significantly larger and have more exons compared with non-epigenes which contributes to increased isoform diversity. Second epigenes participate in more multimeric complexes than non-epigenes. Thirdly, given their proposed importance in membraneless organelles, we show epigenes are enriched for substantially larger intrinsically disordered regions (IDRs). Additionally, we assessed the specificity of their expression profiles and showed epigenes are more ubiquitously expressed consistent with their enrichment in pediatric syndromes with intellectual disability, multiorgan dysfunction, and developmental delay. Finally, in the L1000 dataset, we identify drugs that can potentially be used to modulate expression of these genes.
Conclusions:
Here we identify significant differences in isoform usage, disordered domain content, and variable binding partners between human epigenes and non-epigenes using various functional genomics datasets from Ensembl, ENCODE, GTEx, HPO, LINCS L1000, and BrainSpan. Our results contribute new knowledge to the growing field focused on developing targeted therapies for diseases caused by epigene mutations, such as chromatinopathies and cancers.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
Ligand Binding and Linkage
General Transcription Factors
Cooperative Binding of Transcription Regulators


