蛋白质突变的功能影响可以通过指导选择序列共变信息来预测和解释
Simona Cocco1, Lorenzo Posani1, Rémi Monasson1
1Laboratory of Physics of the Ecole Normale Supérieure, CNRS UMR8023 and Paris Sciences & Lettres (PSL) Research, Sorbonne Université, 75005 Paris, France.
概括
预测蛋白质突变效应是很困难的,因为epistasis. 这种新方法结合了进化和突变数据,以产生高效,可解释的模型,将深度学习性能与更少的参数相匹配.
科学领域:
- 计算生物学 计算生物学
- 蛋白质工程是指蛋白质的工程.
- 生物信息学是一种生物信息学.
背景情况:
- 由于氨基酸之间的复杂相互作用 (epistasis),预测蛋白质突变的功能影响在计算上具有挑战性.
- 现有的方法往往需要大量的数据或缺乏可解释性.
研究的目的:
- 通过整合进化和实验数据,开发一种计算高效的方法来预测突变效应.
- 创建可解释的表观模型,捕捉氨基酸相互作用.
主要方法:
- 一个新的管道,将同源序列数据与有限的突变扫描数据相结合.
- 使用突变发生测量指导在稀疏图形模型中选择链接.
- 从进化序列数据推断模型参数.
主要成果:
- 开发的方法在10个突变扫描上实现了与最先进的深度网络相提并论的性能.
- 这些模型需要更少的参数,提高了可解释性.
- 识别的相互作用是特定于野生类型蛋白质和测量的属性,重点是功能部位.
结论:
- 这种方法提供了一种有效和可解释的方式来预测蛋白质突变效应.
- 它有效地利用进化信息来理解功能约束.
- 该方法提供了有关特定实验环境的蛋白质行为的洞察.
相关概念视频
Mutations
81.8K
Overview
81.8K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Signal Sequences and Sorting Receptors
5.3K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.3K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K


