PDNAPred:基于预先训练的蛋白质语言模型的蛋白质-DNA结合位点的可解释性预测
1College of Information Technology, Shanghai Ocean University, Shanghai 201306, China.
International journal of biological macromolecules
|October 2, 2024
概括
使用一种新的基于序列的方法,PDNAPred准确地识别了蛋白质-DNA结合部位. 该方法的性能优于现有的工具,并显示出预测其他结合地点的潜力.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 蛋白质-DNA相互作用对于生物过程和药物发现至关重要.
- 识别这些相互作用的实验方法耗时且不足以进行大规模分析.
- 现有的计算方法往往需要手动功能或结构数据,限制了它们的适用性.
研究的目的:
- 开发一种高效的,基于序列的计算方法来识别蛋白质-DNA结合点.
- 克服现有方法的局限性,特别是它们依赖于手动功能或结构信息.
- 为了应对在绑定站点预测中不平衡的数据集所带来的挑战.
主要方法:
- 介绍了PDNAPred,一种基于序列的新方法.
- 结合了两个预先训练的蛋白质语言模型和一个定制的CNN-GRU网络.
- 利用焦点损失来处理不平衡的数据集样本.
- 进行模型解释性分析.
主要成果:
- PDNAPred显著提高了DNA结合部位预测的准确性.
- 超越现有的最先进的基于序列的方法.
- 取得的性能与基于结构的先进方法相美.
- 通过成功预测RNA结合部位,证明了多功能性.
结论:
- PDNAPred提供了一个高度准确和高效的解决方案,仅从序列数据中识别蛋白质-DNA结合位点.
- 开发的CNN-GRU网络架构对于绑定站点检测是有效的.
- PDNAPred显示为预测各种氨基酸结合位,包括RNA结合位的可通用框架.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Ligand Binding Sites
12.8K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
12.8K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Protein Organization
6.3K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.3K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K


