通过使用基于MSA的语言模型,提高单个突变引起的蛋白质稳定性变化的预测.
Francesca Cuturello1, Marco Celoria1,2, Alessio Ansuini1
1AREA Science Park, Trieste, 34149, Italy.
Bioinformatics (Oxford, England)
|July 16, 2024
概括
蛋白质语言模型 (PLMs) 可以预测突变导致的蛋白质稳定性变化. 使用进化信息的优化MSA变压器模型显著提高了对现有方法的预测准确性和概括性.
科学领域:
- 结构生物学是结构生物学.
- 计算生物学是一种计算生物学.
- 蛋白质工程是一种蛋白质工程.
背景情况:
- 蛋白质语言模型 (PLM) 在结构生物学中表现有前途,仅使用序列数据.
- 由于有限的数据,预测单氨基酸突变的稳定性变化具有挑战性.
- 现有的方法因数据稀缺和实验限制而扎.
研究的目的:
- 开发一种新的方法来预测由突变引起的蛋白质热力学稳定性转移.
- 在PLM中通过多重序列对齐 (MSA) 利用进化信息.
- 利用大规模数据集与严格的预处理来提高模型性能和防止过拟合.
主要方法:
- 将多个序列对齐 (MSA) 纳入蛋白质语言模型 (PLM).
- 利用了一个大规模的数据集,采用严格的数据预处理和数据泄露防治政策.
- 对各种精心调整的预训练模型进行了比较分析,包括废除研究和基线评估.
主要成果:
- 通过利用来自MSA的共同进化信号,MSA变压器表现出卓越的准确性.
- 优化的MSA变压器在预测蛋白质稳定性变化方面表现优于现有的方法.
- 该模型表现出增强的概括能力,改善了对点突变的预测.
结论:
- 通过MSA集成进化信息可以显著提高稳定性预测的PLM性能.
- 优化的MSA变压器代表了一种先进的方法,用于预测突变诱导的稳定性变化.
- 这种方法为蛋白质工程和理解蛋白质功能提供了一个强大的工具.
更多相关视频
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
7.3K
05:56Exploring Caspase Mutations and Post-Translational Modification by Molecular Modeling Approaches
Published on: October 13, 2022
1.3K
相关概念视频
RNA Stability
33.5K
Intact DNA strands can be found in fossils, while scientists sometimes struggle to keep RNA intact under laboratory conditions. The structural variations between RNA and DNA underlie the differences in their stability and longevity. Because DNA is double-stranded, it is inherently more stable. The single-stranded structure of RNA is less stable but also more flexible and can form weak internal bonds. Additionally, most RNAs in the cell are relatively short, while DNA can be up to 250 million...
33.5K
mRNA Stability and Gene Expression
2.8K
2.8K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
