用基于序列的蛋白质语言模型绘制蛋白质结合点的空间
Tuğçe Oruç1, Maria Kadukova1, Thomas G Davies1
1Computational Chemistry & Informatics, Astex Pharmaceuticals, Cambridge CB4 0QA, United Kingdom.
Bioinformatics (Oxford, England)
|June 27, 2025
概括
我们开发了一种新方法,将蛋白质语言模型和3D图形集成在一起,以表示蛋白质结合点. 这种方法通过改善结合部位的比较和上下文化来增强药物发现.
科学领域:
- 计算生物学是一种计算生物学.
- 结构生物信息学 结构生物信息学
- 药物发现 药物发现
背景情况:
- 蛋白质结合部位对生物活动至关重要,是治疗干预的关键目标.
- 检测,比较和语境结合点的有效方法在药物发现中具有重要意义.
研究的目的:
- 提出一种用于生成蛋白质结合位点详细表示的新方法.
- 为了证明这些表示在药物发现下游机器学习任务中的实用性.
主要方法:
- 蛋白质语言模型与3D图形化技术的整合.
- 结合功能,结构和进化信息的绑定站点表示的衍生.
- 开发相似度指标,平衡本地结构和全球序列效应.
主要成果:
- 创建了丰富多功能,具有前所未有的细节的绑定站点表示.
- 证明相似度指标诱导有意义的口袋集群.
- 展示了下游任务的简化,包括口袋组组织,药物适应性建模和基准测试.
结论:
- 开发的方法为理解和利用蛋白质结合部位信息提供了强大的嵌入.
- 这种方法促进了新结合部位的高效上下文化,并改善了用于药物发现的机器学习模型的开发.
相关概念视频
Conserved Binding Sites
4.4K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.4K
Ligand Binding Sites
13.4K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
13.4K
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Protein-protein Interfaces
13.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
13.5K
Ligand Binding and Linkage
4.9K
Allosteric proteins have more than one ligand binding site; the binding of a ligand to any of these sites influences the binding of ligands to the other sites. When a protein is allosteric, its binding sites are called coupled or linked. In the case of enzymes, the site that binds to the substrate is known as the active site and the other site is known as the regulatory site. When a ligand binds to the regulatory site, this leads to conformational changes in the protein that can influence...
4.9K
Protein Families
15.9K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.9K


