残留物保存和溶剂可访问性是 (几乎) 所有你需要预测蛋白质的突变效应
Matsvei Tsishyn1,2, Pauline Hermans1,2, Marianne Rooman1,2
1Computational Biology and Bioinformatics, Université Libre de Bruxelles, Brussels 1050, Belgium.
Bioinformatics (Oxford, England)
|May 28, 2025
概括
一个简单的进化得分优于复杂的深度学习模型在预测蛋白质特性上的突变效应方面. 这一发现突显了当前计算生物学预测器评估突变景观的局限性.
科学领域:
- 计算生物学是一种计算生物学.
- 蛋白质工程是一种蛋白质工程.
- 生物信息学是一种生物信息学.
背景情况:
- 预测突变对蛋白质生物物理性质的影响是一个重大挑战.
- 现有的深度学习模型往往缺乏解释性,准确性有限.
研究的目的:
- 开发和评估一种新的,可解释的模型来预测突变影响.
- 为了比较一个简单的进化得分与复杂的预测器的表现.
主要方法:
- 开发了RSALOR模型,使用基于残留频率和溶剂可访问性的进化得分.
- 在ProteinGym数据集上评估RSALOR,评估稳定性,活性和适应性.
- 基准RSALOR与各种复杂的预测指标相比.
主要成果:
- RSALOR的表现与大多数复杂的基准预测指标相提并论,甚至比它们更好.
- 一个简单的进化得分可以在各种突变数据集上实现高预测性能.
- 该研究质疑复杂的深度学习模型的学习能力和局限性.
结论:
- 简单的,可解释的模型可以在预测突变效应方面非常有效.
- 强调需要重新评估计算生物学中复杂的深度学习模型的实用性和可解释性.
- RSALOR提供了一个用户友好的Python包,用于预测突变影响.
相关概念视频
Conserved Binding Sites
4.4K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.4K
Chemical Shift: Internal References and Solvent Effects
826
In an NMR sample, precise measurement of the absolute absorption frequencies of nuclei is difficult. A standard internal reference compound is added, and the frequency difference between the reference signal and sample signals is measured.
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...
826
Mismatch Repair
5.2K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
5.2K
Covalently Linked Protein Regulators
7.5K
Proteins can undergo many types of post-translational modifications, often in response to changes in their environment. These modifications play an important role in the function and stability of these proteins. Covalently linked molecules include functional groups, such as methyl, acetyl, and phosphate groups, and also small proteins, such as ubiquitin. There are around 200 different types of covalent regulators that have been identified.
These groups modify specific amino acids in a protein....
These groups modify specific amino acids in a protein....
7.5K
Conservation of Protein Domains Over Different Proteins
11.5K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.5K
Spontaneous and Induced Mutations
181
Spontaneous mutations arise infrequently during DNA replication due to errors in the process. A key factor behind these errors is tautomeric shifts in nitrogenous bases, where bases transition from keto to enol forms or amino to imino forms. This shift can alter base-pairing rules, leading to mutations. Additionally, reactive oxygen species (ROS) arising from aerobic metabolism can damage DNA, resulting in depurination (loss of a purine base) or depyrimidination (loss of a pyrimidine base).
181


