数据效率高的蛋白质突变效应预测,通过分子模拟和蛋白质语言模型的弱监督
Teppei Deguchi1,2, Nur Syatila Ab Ghani3, Yoichi Kurumida3
1Graduate School of Frontier Sciences, The University of Tokyo, 5-1-5, Kashiwanoha, Kashiwa, Chiba 277-0882, Japan.
Briefings in bioinformatics
|October 10, 2025
概括
这项研究引入了一种新的数据增强方法,用于预测蛋白质突变的机器学习模型. 它提高了蛋白质工程和病原性分析的预测准确性,特别是在有限的实验数据下.
科学领域:
- 计算生物学 计算生物学
- 生物物理学的生物物理.
- 机器学习 机器学习
背景情况:
- 预测蛋白质突变效应对于蛋白质工程和病原性评估至关重要.
- 由于实验数据有限和成本高昂,目前的方法面临挑战.
- 以前的数据增强依赖于分子模拟,仅限于热稳定性.
研究的目的:
- 开发一种用于蛋白质突变效应预测的新数据增强技术.
- 将计算预测的适用性扩展到各种蛋白质特性.
- 为了提高预测准确性在低数据的制度.
主要方法:
- 从蛋白质语言模型进行数据增强的零射击预测的组合分子模拟.
- 使用计算估计作为"弱"的训练数据.
- 根据实验数据的可用性,动态调整弱数据的权重和包含.
主要成果:
- 新方法提高了预测的准确性,特别是当实验数据稀缺时.
- 成功扩展了对蛋白质结合亲和力和酶活性预测的适用性.
- 在小数据场景的基准测试中表现有所改善.
结论:
- 提出的方法有效地补充了机器学习模型的实验数据.
- 提供了一个强大的方法,用于蛋白质工程和病原性预测有限的数据.
- 在数据稀缺的环境中推进蛋白质突变效应预测领域.
更多相关视频
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
2.5K
11:36A Protocol for Functional Assessment of Whole-Protein Saturation Mutagenesis Libraries Utilizing High-Throughput Sequencing
Published on: July 3, 2016
11.3K
相关概念视频
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Covalently Linked Protein Regulators
8.7K
Proteins can undergo many types of post-translational modifications, often in response to changes in their environment. These modifications play an important role in the function and stability of these proteins. Covalently linked molecules include functional groups, such as methyl, acetyl, and phosphate groups, and also small proteins, such as ubiquitin. There are around 200 different types of covalent regulators that have been identified.
These groups modify specific amino acids in a protein....
These groups modify specific amino acids in a protein....
8.7K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
