探索进化以揭示对蛋白质突变稳定性的洞察力
Pauline Hermans1,2, Matsvei Tsishyn1,2, Martin Schwersensky1,2
1Computational Biology and Bioinformatics, Université Libre de Bruxelles, Brussels 1050, Belgium.
Molecular biology and evolution
|January 9, 2025
概括
预测突变导致的蛋白质稳定性变化至关重要. 令人惊的是,简单的进化模型与残留物可访问性相匹配的复杂方法,为蛋白质设计和变体解释提供了新的见解.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 进化生物学 进化生物学
背景情况:
- 蛋白质的热力学稳定性对于蛋白质的设计和解释遗传变异至关重要.
- 从同源蛋白序列中获得的进化数据通常用于预测突变对稳定性的影响.
- 深度突变扫描提供了广泛的数据来完善这些预测方法.
研究的目的:
- 研究构建多个序列对齐的最佳方法,并提取用于蛋白质稳定性预测的进化信息.
- 评估不同进化模型的有效性,包括独立地点和表观模型.
- 评估结构特征的贡献,如溶剂可访问性,结合进化数据.
主要方法:
- 利用大规模的深度突变扫描稳定性数据.
- 构建和分析各种多重序列对齐.
- 测试了独立地点和复杂的表观进化模型.
- 作为结构特征,纳入了突变残留物的相对溶剂可访问性.
主要成果:
- 独立站点进化模型的准确性与更复杂的表观模型可比,用于稳定性预测.
- 复杂的表观模型通常会产生噪音合,而不会比更简单的模型有显著的预测改进.
- 将进化特征与相对溶剂可访问性相结合,实现了类似于先进机器学习预测器的预测准确性.
结论:
- 简单的进化模型对于预测突变诱导的稳定性变化是有效的.
- 相对溶剂可访问性是一个有价值的特征,当与进化数据相结合时,可以提高预测准确性.
- 这些发现为利用进化信息来改善蛋白质稳定性预测提供了新的视角.
更多相关视频
相关概念视频
Gene Evolution - Fast or Slow?
7.0K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.0K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Mismatch Repair
4.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
4.8K
Covalently Linked Protein Regulators
6.8K
Proteins can undergo many types of post-translational modifications, often in response to changes in their environment. These modifications play an important role in the function and stability of these proteins. Covalently linked molecules include functional groups, such as methyl, acetyl, and phosphate groups, and also small proteins, such as ubiquitin. There are around 200 different types of covalent regulators that have been identified.
These groups modify specific amino acids in a protein....
These groups modify specific amino acids in a protein....
6.8K
Mutations
79.5K
Overview
79.5K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


