基于机器学习的蛋白质结构预测,使用氨基酸序列和结构字母表
Jad Abbass1, Charles Parisi1,2
1School of Computer Science and Mathematics, Kingston University, London, UK.
Journal of biomolecular structure & dynamics
|March 20, 2024
概括
一个新的机器学习模型使用氨基酸序列自动将蛋白质域分类为CATH架构. 这种方法有助于注释庞大的蛋白质结构数据库和预注释序列,改善结构生物学研究.
科学领域:
- 结构生物学 结构生物学
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
背景情况:
- 蛋白质数据库 (PDB) 和AlphaFold预测已经大大增加了蛋白质结构数据.
- 使用CATH域数据库对这些结构进行分类至关重要,但由于手动注释的限制,具有挑战性.
- 下一代测序产生了许多缺乏结构注释的蛋白质序列,需要有效的预注释方法.
研究的目的:
- 开发一个完全自动化的机器学习模型,用于将蛋白质域分类为CATH架构.
- 为了应对快速扩展的蛋白质结构存储库的注释挑战.
- 为了能够准确地预先注释具有结构特征的蛋白质序列.
主要方法:
- 开发了一种用于蛋白质域分类的新型机器学习模型.
- 使用氨基酸序列作为模型的输入.
- 嵌入的结构字母表与氨基酸序列一起用于增强分类.
主要成果:
- 仅使用氨基酸序列,获得了0.92的F1得分.
- 在使用氨基酸序列和结构字母表时,获得了0.94的F1评分.
- 证明了模型对已知和未知的蛋白质结构进行分类的能力.
结论:
- 开发的机器学习模型为CATH架构分类提供了高度准确和自动化的解决方案.
- 这种方法可以在注释大型蛋白质结构数据集和预注释新序列方面显著帮助.
- 该模型为生物信息学中的结构和功能注释提供了有价值的工具.
关键词:
在CATH系统中,CATH目的范围 SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP SCOP S这就是K-MER.机器学习是机器学习.蛋白质块可以阻断蛋白质.蛋白质的架构是蛋白质的结构.蛋白质的结构 蛋白质的结构结构字母表结构字母表相关概念视频
Protein Organization
137.6K
Overview
137.6K
Protein Folding
118.1K
Overview
118.1K
Protein and Protein Structure
79.5K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
79.5K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Protein and Protein Structures
10.4K
10.4K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K


