输入序列对共识衍生的蛋白质的重要性及其与重建的祖先蛋白质的关系
Charlotte Nixon1, Shion A Lim1, Matt Sternke2
1Department of Molecular and Cell Biology, University of California, Berkeley, Berkeley, California, USA.
概括
遗传学分析显示,来自不同蛋白家族的共识蛋白在从密切相关的序列中衍生时更稳定. 这表明,在广泛的共识产生过程中,蛋白质合作性失去了关节特异性机制.
科学领域:
- 蛋白质的生物信息学
- 分子进化是分子进化的过程.
- 结构生物学是结构生物学.
背景情况:
- 蛋白质序列决定了它的能量格局,包括构造,能量和动态.
- 对同源序列的基因分析可以揭示进化关系,并有助于重建祖先或共识蛋白.
- 祖先和共识蛋白都通常比现有的同类更稳定,这表明工程热稳定性的潜力.
研究的目的:
- 为了比较祖先序列重建 (ASR) 和工程热稳定性共识蛋白质生成方法.
- 评估输入序列的进化相关性如何影响使用Ribonuclease H家族产生的共识蛋白质的特性.
- 研究与序列多样性相关的蛋白质稳定性和合作性的机制.
主要方法:
- 编译了Ribonuclease H.家族的多个序列对齐.
- 从完整和遗传学上受限制的序列对齐中生成共识蛋白.
- 评估蛋白质特性,包括结构,活性,稳定性和合作折叠.
- 分析了使用波特形式主义 (对对共变) 和单数值分解 (SVD) 的序列相关性.
主要成果:
- 来自完全Ribonuclease H对齐的共识蛋白是活跃的,但缺乏增强的稳定性和合作折叠.
- 来自基因组学上受限制的对齐的共识蛋白显示出显著增强的稳定性和合作折叠.
- SVD分析显示,稳定的共识序列与其祖先和后代聚集在一起,而不稳定的则是异常值.
- 稳定和不稳定的共识蛋白之间的对对共变性和更高阶相关性不同.
结论:
- 序列的进化相关性显著影响共识蛋白的稳定性和折叠合作性.
- 从族系遗传学上受限的集合中生成共识蛋白,可以保留类特异性的合作机制,从而提高稳定性.
- 通过SVD分析,通过它们在序列空间中的位置与进化轨迹相对,有效地区分稳定和不稳定的共识蛋白.
相关概念视频
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K


