CNNCaps-DBP:利用注意力增强卷积的蛋白质语言模型进行DNA结合蛋白质预测
Ziyuan Yan1, Aoyun Geng1, Yazi Li2
1School of Computer Science and Technology, Hainan University, Haikou, 570228, China.
概括
一种新的深度学习方法,CNNCaps-DBP,使用序列信息准确预测DNA结合蛋白 (DBPs). 这种计算方法超越了现有的模型,提供了一种更快,更有效的方法来识别DBP,这对于了解细胞过程和疾病至关重要.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 结合DNA的蛋白质 (DBPs) 对于DNA复制和基因调节至关重要,在癌症等疾病中发挥作用.
- 实验性DBP识别是耗时且昂贵的.
- 当前的预测模型往往无法有效利用预训练的蛋白质语言模型的特征.
研究的目的:
- 开发一种新,准确和高效的计算方法,仅使用初级序列信息来预测DNA结合蛋白 (DBPs).
- 解决现有的DBP预测模型的局限性,特别是从预训练模型中提取它们的特征.
主要方法:
- 提出CNNCaps-DBP,这是一个深度学习模型,集成ESM C预训练的蛋白质语言模型.
- 采用注意力增强的卷积模块来增强蛋白质嵌入.
- 利用混合囊网络和MLP架构进行预测,并通过动态学习速率调度器进行优化.
主要成果:
- 与现有模型相比,CNNCaps-DBP的预测性能明显优于现有模型.
- 该模型在独立数据集上保持了高性能,超过了最先进的方法.
- 案例研究证实了该模型在DBP识别方面的强大预测能力.
结论:
- CNNCaps-DBP提供了一个强大而高效的计算框架,用于从序列数据中准确预测DBP.
- 该方法克服了以前方法的局限性,通过有效利用预先训练的蛋白质语言模型.
- 这一进步有助于研究与DBP相关的蛋白质功能和疾病机制.
相关概念视频
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Ligand Binding Sites
14.9K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
14.9K
Protein-protein Interfaces
14.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.4K
Single-Strand DNA Binding Proteins
16.5K
For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
16.5K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
From DNA to Protein
21.9K
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
21.9K


