用深度卷积神经网络对基因表达和DNA结合部位的定位进行预测建模
Arman Karshenas1, Tom Röschinger2, Hernan G Garcia1,3,4,5,6
1Biophysics Graduate Group, University of California at Berkeley, Berkeley, CA, USA.
bioRxiv : the preprint server for biology
|January 7, 2025
概括
深度学习可适应的调控序列标识符 (DARSI) 使用MPRA数据来预测基因表达和识别转录因子结合部位. 这个工具通过揭示调节性DNA序列功能来增强基因组注释.
科学领域:
- 基因组学就是基因组学.
- 计算生物学 计算生物学
- 分子生物学分子生物学
背景情况:
- 基因组测序已经进步,但对调节性DNA上的转录因子结合部位的排列仍然在很大程度上没有特征.
- 大规模并行记者测试 (MPRA) 提供了一种方法来测量由数千种调节性DNA变异驱动的基因表达,但它们的分析通常假定独立的基对贡献.
- 目前的方法很难解释调节序列中远距离基体之间的相关性,限制了全面的基因组注释.
研究的目的:
- 开发一种计算工具,通过考虑监管DNA中远距离基之间的相关性来分析大规模并行报告测试 (MPRA) 数据.
- 为了能够直接从DNA序列中准确预测基因表达水平.
- 在监管区域内系统地确定单基对分辨率的转录因子结合点.
主要方法:
- 开发了深度学习适应性调节序列标识符 (DARSI),一个卷积神经网络.
- 训练有素的DARSI使用MPRA数据从原始调节性DNA序列预测基因表达.
- 通过对已知转录因子结合部位的精心策划的数据库进行比较,验证了DARSI的预测.
主要成果:
- DARSI准确地预测了转录因子结合部位,证明了与已建立的数据库的高度一致性.
- 该模型成功地确定了新的,以前未被绘制的转录因子结合部位.
- 达西的预测为未来的研究提供了实验性可操作的见解.
结论:
- 通过自动识别转录因子结合位点,DARSI提高了调节性DNA区域的注释.
- 该工具预测结合点的能力有助于理解转录控制机制.
- 达西的预测指导了未来的实验验证,推动了基因组学中的理论-实验循环.
相关概念视频
Chromatin Position Affects Gene Expression
23.2K
Chromatin is the massive complex of DNA and proteins packaged inside the nucleus. The complexity of chromatin folding and how it is packaged inside the nucleus greatly influences access to genetic information. Generally, the nucleus' periphery is considered transcriptionally repressive, while the cell's interior is considered a transcriptionally active area.
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
Topologically Associated Domains (TADs)
The 3-dimensional positioning of chromatin in the nucleus influences the...
23.2K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
DNA Microarrays
17.2K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.2K
Reporter Genes
11.2K
Reporter genes are a type of protein-coding gene that are often tagged to a gene of interest. Once inside a target cell, reporter genes usually produce visually identifiable characteristics like fluorescence and luminescence when expressed along with the gene of interest. Thus, reporter genes “report” the presence or absence of genes of interest in an organism, determine the gene expression pattern, or track the physical location of a DNA segment or protein in the cell.
11.2K


