Deep learning sequence-based ab initio prediction of variant effects on expression and disease risk
Jian Zhou1,2,3, Chandra L Theesfeld1, Kevin Yao3
1Lewis-Sigler Institute for Integrative Genomics, Princeton University, Princeton, NJ, USA.
Nature Genetics
|July 18, 2018
Summary
We created ExPecto, a deep learning tool that predicts how DNA mutations affect gene expression in specific tissues. This framework aids in understanding genetic diseases and evolutionary biology by analyzing vast amounts of genomic data.
Area of Science:
- Genomics
- Computational Biology
- Systems Biology
Background:
- Deciphering gene expression regulation and the impact of genomic variations is crucial for human genetics, precision medicine, and evolutionary biology.
- The vast scale of the noncoding genome presents a significant challenge in understanding mutation effects.
Purpose of the Study:
- To develop a deep learning framework, ExPecto, for accurate *ab initio* prediction of tissue-specific transcriptional effects of DNA mutations.
- To analyze the regulatory mutation space and its implications for gene expression and disease risk.
Main Methods:
- Developed ExPecto, a deep learning framework predicting mutation effects on gene expression from DNA sequence.
- Applied ExPecto to prioritize causal variants in genome-wide association studies (GWAS) loci.
- Conducted *in silico* saturation mutagenesis to characterize regulatory mutation space for over 140 million promoter-proximal mutations.
Main Results:
- ExPecto accurately predicts *ab initio* tissue-specific transcriptional effects of mutations, including rare and unobserved ones.
- Experimental validation confirmed predictions for four immune-related diseases.
- Characterized the regulatory mutation space for human RNA polymerase II-transcribed genes, enabling analysis of evolutionary constraints.
Conclusions:
- ExPecto provides an end-to-end computational framework for predicting gene expression and disease risk from DNA sequence.
- The framework facilitates the study of evolutionary constraints on gene expression and the prediction of mutation-induced disease effects.
Related Concept Videos
Histone Variants at the Centromere
5.1K
Histone variants are the histone proteins with structural and sequence variations. These variants may be regarded as “mutant” forms that replace their canonical histone counterparts in the nucleosomes. Specific post-translational modifications on the histone variants enable further chromatin complexity and regulate tissue-specific gene expression. The most common histone variants are from histone H2A, H2B, and linker histone H1 families. However, several variants of histone H3...
5.1K
Predicting Molecular Geometry
46.0K
VSEPR Theory for Determination of Electron Pair Geometries
46.0K
Relative Risk
2.2K
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
2.2K
Cis-regulatory Sequences
11.9K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.9K
Prediction Intervals
3.4K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.4K
What is Gene Expression?
196.9K
Overview
Gene expression is the process in which DNA directs the synthesis of functional products, that is, proteins. Cells can regulate gene expression at various stages. It allows organisms to generate different cell types and enables cells to adapt to internal and external factors.
Genetic Information Flows from DNA to RNA to Protein
A gene is a stretch of DNA that serves as the blueprint for functional RNAs and proteins. Since DNA is made up of nucleotides and proteins consist of amino...
Gene expression is the process in which DNA directs the synthesis of functional products, that is, proteins. Cells can regulate gene expression at various stages. It allows organisms to generate different cell types and enables cells to adapt to internal and external factors.
Genetic Information Flows from DNA to RNA to Protein
A gene is a stretch of DNA that serves as the blueprint for functional RNAs and proteins. Since DNA is made up of nucleotides and proteins consist of amino...
196.9K


