Cell type-specific interpretation of noncoding variants using deep learning-based methods
Maria Sindeeva1, Nikolay Chekanov1, Manvel Avetisian1
1AIRI, Moscow, 121170, Russia.
Gigascience
|March 27, 2023
Summary
DeepCT, a novel neural network, predicts noncoding variant effects across cell types by inferring missing epigenetic data. This advances human genetics by overcoming data limitations in machine learning models.
Area of Science:
- Genomics
- Computational Biology
- Machine Learning
Background:
- Interpreting noncoding genomic variants is a major challenge in human genetics.
- Machine learning (ML) models can predict the effects of noncoding mutations but require extensive, cell-type-specific experimental data.
- Existing ML approaches are limited by the sparsity of available epigenetic data across human cell types.
Purpose of the Study:
- To develop a novel ML approach that overcomes data limitations for predicting noncoding variant effects.
- To infer missing epigenetic features and generalize predictions across diverse cell types.
- To enable cell type-specific predictions of noncoding variant impacts.
Main Methods:
- Proposed DeepCT, a new neural network architecture designed to learn interconnections within epigenetic features.
- Developed DeepCT to infer unmeasured epigenetic data from available inputs.
- Enabled DeepCT to learn cell type-specific properties and generate biologically meaningful vector representations of cell types.
Main Results:
- DeepCT successfully infers missing epigenetic data, addressing the sparsity issue.
- The model learns cell type-specific characteristics, creating distinct vector representations.
- DeepCT generates accurate cell type-specific predictions for the effects of noncoding genomic variations.
Conclusions:
- DeepCT offers a powerful solution for interpreting noncoding variants by overcoming epigenetic data scarcity.
- The ability to infer data and generalize across cell types significantly advances the application of ML in human genetics.
- This approach facilitates more precise, cell type-specific predictions of variant effects, aiding genetic research and clinical applications.
More Related Videos
Related Concept Videos
lncRNA - Long Non-coding RNAs
8.7K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.7K
Comparing Copy Number Variations and SNPs
17.8K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.8K
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K


