基于深度学习的cis调节性DNA序列的语义匹配有助于预测基因功能
Tianyi Li1,2,3, Hui Xu1, Mingrui Suo1
1State Key Laboratory of Maize Bio-breeding, Frontiers Science Center for Molecular Design Breeding, Joint International Research Laboratory of Crop Molecular Breeding, National Maize Improvement Center, College of Agronomy and Biotechnology, China Agricultural University, Beijing, People's Republic of China.
Nature plants
|February 18, 2026
概括
一个新的深度学习模型,PhytoBabel,捕捉了与远距离相关的植物cis-regulatory DNA 序列中的语义相似性. 这通过发现新的基因关系,推进了逆转基因学中的基因功能预测.
科学领域:
- 基因组学就是基因组学.
- 计算生物学 计算生物学
- 植物科学 植物科学
背景情况:
- 基因调节DNA序列包含对基因功能预测至关重要的信息.
- 在反向遗传学中利用这些信息仍然具有挑战性,因为序列分歧.
研究的目的:
- 开发一个深度学习模型 (PhytoBabel),以捕捉正义 cis-regulatory 序列中的语义相似性.
- 为了证明PhytoBabel对基因功能预测和发现新型基因关系的实用性.
主要方法:
- 培训PhytoBabel对来自15种种子的正交管调节序列进行正交管调节.
- 使用深度学习识别语义相似性,尽管低序列同源性.
- 通过识别与已知的Arabidopsis调节者语义上相似的玉米基因来验证模型.
主要成果:
- 菲托巴贝尔有效地捕捉了在1亿6000万年间分离的cis-regulatory序列中的语义相似性.
- 该模型隐式学习了基因表达模式,保存序列和家族遗传关系.
- 确定了进化上无关但语义上相似的 cis-regulatory 序列,有助于新基因的发现.
结论:
- 植物宝贝弥合了 cis 调节序列,语义和基因功能之间的差距.
- 该模型为逆转基因学中的基因功能预测提供了一个强大的新工具.
- 促进发现具有跨物种保留功能的新基因.
相关概念视频
Cis-regulatory Sequences
12.0K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
12.0K
Cis-regulatory Sequences
4.2K
4.2K
Cooperative Binding of Transcription Regulators
7.4K
Transcriptional regulators bind to specific cis-regulatory sequences in the DNA to regulate gene transcription. These cis-regulatory sequences are very short, usually less than ten nucleotide pairs in length. The short length means that there is a high probability of the exact same sequence randomly occurring throughout the genome. Since regulators can also bind to groups of similar sequences, this further increases the chances of random binding. Transcriptional regulators form...
7.4K
Conserved Binding Sites
5.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.2K
Co-activators and Co-repressors
8.7K
Gene transcription is regulated by the synergistic action of several proteins that form a complex at a gene regulatory site. This is observed in eukaryotes, where the regulation of gene expression is a complex process. Regulatory proteins in eukaryotes can broadly be classified into two types – regulators that bind directly to specific DNA sequences and co-regulators that associate with regulatory proteins but cannot directly bind to the DNA. These co-regulators are further divided into...
8.7K
Master Transcription Regulators
7.9K
Master transcription regulators are regulatory proteins that are predominantly responsible for regulating the expression of multiple genes. Often these genes work in concert to drive a complex process. Activation of a master transcription regulator can lead to a cascade of transcriptional activation necessary for that outcome. These regulators can directly bind to the regulatory sequences of the various genes involved, or they can indirectly regulate transcription by binding to regulatory...
7.9K


