通过对DNA结合蛋白的域适应性预训,提高一般蛋白语言模型的预测性能
Wenwu Zeng1, Yutao Dou1, Liangrui Pan1
1College of Computer Science and Electronic Engineering, Hunan University, Changsha, 410082, China.
Nature communications
|September 7, 2024
概括
我们开发了ESM-DBP,这是一种新的计算方法,通过精炼蛋白质语言模型与特定领域的知识来改进DNA-蛋白质相互作用的识别. 这提高了对关键生物过程的预测准确度.
科学领域:
- 计算生物学 计算生物学
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
背景情况:
- DNA-蛋白相互作用对于DNA复制,转录和基因调节至关重要.
- 目前用于识别DNA-蛋白相互作用的计算方法缺乏准确性和效率.
- 一般的蛋白质语言模型还没有完全探索DNA结合蛋白质域特定的知识.
研究的目的:
- 提出一种新的计算方法,ESM-DBP,用于准确有效地识别DNA-蛋白相互作用.
- 利用精细的DNA结合蛋白序列数据集,利用域适应性预训练.
- 为了改善DNA结合蛋白的特征表示.
主要方法:
- 通过选170,264个DNA结合蛋白序列,构建了一个域适应性语言模型.
- 改进了用于预训练的DNA结合蛋白序列库.
- 在四个下游任务上评估ESM-DBP并与现有方法进行比较.
主要成果:
- 与原始语言模型相比,ESM-DBP为DNA结合蛋白提供了优越的特征表示.
- 该方法显著提高了预测性能,超过了最先进的技术.
- 即使在具有有限同源序列的序列上,ESM-DBP也表现出强大的性能.
- 在两种情况下,ChIP-seq实验验证了该方法的预测.
结论:
- ESM-DBP提供了一种更有效的方法来识别DNA-蛋白相互作用.
- 域自适应预训练策略增强了模型捕获特定生物知识的能力.
- 这种方法推进了研究基因调节和其他DNA-蛋白质介导过程的计算工具.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Conservation of Protein Domains
3.1K
3.1K
Ligand Binding Sites
12.8K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
12.8K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Improving Translational Accuracy
9.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.4K


