通过多种蛋白质语言模型和合奏学习模型准确和快速预测内在无序的蛋白质
Shijie Xu1, Akira Onoda1,2
1Graduate School of Environmental Science, Hokkaido University, Sapporo 060-0810, Japan.
Journal of chemical information and modeling
|October 26, 2023
概括
我们开发了IDP-ELM,这是一种使用蛋白质语言模型的新方法,仅从蛋白质序列中准确预测内在无序区域 (IDR) 和它们的功能. 这种快速方便的工具有助于蛋白质层面的分析.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
背景情况:
- 内在无序蛋白 (IDP) 对于生物过程至关重要,并且从初级序列预测它们对于蛋白质分析至关重要.
- 机器学习,特别是蛋白质语言模型 (PLM),显示出对高效和精确的蛋白质序列分析有很大的希望.
研究的目的:
- 开发一种新的方法,IDP-ELM,用于预测内在无序区域 (IDR) 和它们的功能,如灵活的链接器和蛋白质结合点.
- 利用先进的PLM和集体学习来提高IDP预测的准确性.
主要方法:
- 利用了来自最先进的PLM的高维表示.
- 用双向循环神经网络进行IDR预测.
- 在独立的CAID和CAID2数据集上评估性能.
主要成果:
- IDP-ELM在曲线下面积 (AUC),马修相关系数 (MCC) 和F1得分方面表现出显著的改善.
- 该方法只需要蛋白质序列,因此无需耗时生成资料.
- 实现了准确,快速,方便的蛋白质层次分析.
结论:
- IDP-ELM为预测IDR及其功能提供了一个强大而有效的工具.
- 该方法的仅序列输入和高性能使其对大规模蛋白质组学研究具有价值.
- 可复制的代码和模型重量是公开可用的,以便进一步研究.
相关概念视频
Intrinsically Disordered Proteins
17.9K
Intrinsically disordered proteins are a group of proteins that do not fold into specific three-dimensional structures. Their structural flexibility allows them to complement ordered proteins to perform functions that are inaccessible to rigid structures. They are more common in eukaryotes than prokaryotes and may either be exclusively intrinsically disordered or hybrid proteins, consisting of a mix of ordered and disordered regions. The absence of a rigid structure in these proteins can be...
17.9K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Improving Translational Accuracy
11.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.4K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


