结还是没有结? 用基于序列的机器学习模型识别结节家族中的未结节蛋白.
Maciej Sikora1,2, Eva Klimentova3,4, Dawid Uchal1,5
1Centre of New Technologies, University of Warsaw, Warsaw, Poland.
Protein science : a publication of the Protein Society
|June 18, 2024
概括
这项研究使用AlphaFold预测来识别UniProt数据库中的结结蛋白质. 一个机器学习模型从序列中预测蛋白质结,揭示了大多数家族中保存的结结构.
科学领域:
- 结构生物学是结构生物学.
- 计算生物学是一种计算生物学.
- 机器学习是机器学习.
背景情况:
- 结结的蛋白质很少见,但在结构上是至关重要的,目前正在研究它们的功能.
- AlphaFold提供了广泛的蛋白质结构预测,使大规模的计算研究成为可能.
研究的目的:
- 在UniProt数据库中使用AlphaFold预测计算识别和分析结节和未结节的蛋白质.
- 开发一种机器学习模型,从氨基酸序列中预测蛋白结的存在.
- 为了研究蛋白质家族中节点的结构和进化保护.
主要方法:
- 利用AlphaFold的预测,构建了来自UniProt的结结和未结结的蛋白质的综合数据集.
- 开发并验证了一种基于序列的结预测的机器学习模型.
- 将结结的蛋白质分类为结构家族,并分析了序列的保存.
主要成果:
- 成功创建了一个结节和未结节蛋白质的强大数据集.
- 在基于序列的结预测100种蛋白质的测试组上达到了92%的一致性.
- 确定所有预测的结结蛋白只属于17个家族.
- 发现了三个新的家族 (UCH,DUF4253,DUF2254),其中有结结的和没有结结的成员,可能是由于删除.
- 发现结结的拓在15个家族中的11个家族中被保存,即使序列相似性很低.
- 确定预测的未结结的蛋白质在结构上是准确的,但可能是非功能碎片.
结论:
- 机器学习模型可以从氨基酸序列准确地预测蛋白质结结.
- 结结的蛋白质结构在特定的蛋白质家族中高度保存.
- 由AlphaFold预测的未结合的蛋白质可能代表截断或非功能变体.
相关概念视频
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Conservation of Protein Domains
3.1K
3.1K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


