输入序列对共识衍生的蛋白质的重要性及其与重建的祖先蛋白质的关系
Charlotte Nixon1, Shion A Lim1, Matt Sternke2
1Department of Molecular and Cell Biology, University of California, Berkeley, Berkeley, CA 94720.
bioRxiv : the preprint server for biology
|July 10, 2023
概括
祖先序列重建和共识蛋白可以设计热稳定性. 然而,来自不同序列的共识蛋白缺乏稳定性,与来自受限制类的蛋白质不同,这表明了类特异性蛋白质合作机制.
科学领域:
- 蛋白质的生物信息学
- 结构生物学是结构生物学.
- 进化生物学是进化的生物学.
背景情况:
- 蛋白质序列决定了它的能量格局,包括结构,能量和动力学.
- 对同源序列的遗传学分析可以揭示序列和景观之间的进化关系.
- 祖先和共识蛋白通常比现有的同类更稳定,这表明热稳定性工程的潜力.
研究的目的:
- 为了比较祖先序列重建和共识蛋白质对工程热稳定性的方法.
- 评估输入序列的进化关系如何影响共识蛋白质属性.
- 研究蛋白质合作性和稳定性背后的机制.
主要方法:
- 使用Ribonuclease H蛋白家族进行比较分析.
- 从广泛的和被遗传学限制的序列集合中生成共识蛋白.
- 采用波特的形式主义,用于对对的协差分析.
- 应用单数值分解 (SVD) 来分析高阶合.
主要成果:
- 总体共识的蛋白质有结构和活性,但缺乏增强的稳定性和合作折叠.
- 来自遗传学上受限制地区的共识蛋白显示出显著增强的稳定性和合作折叠.
- SVD分析显示,稳定的共识序列与其祖先和后代聚集在一起,而不稳定的则是异常值.
结论:
- 蛋白质的合作性可能被不同的进化类别内不同的机制编码.
- 在共识蛋白生成中结合过多的多样性可以导致合作机制和稳定性的丧失.
- 基于其进化背景,SVD分析有效地区分稳定与不稳定的共识蛋白.
相关概念视频
Multi-species Conserved Sequences
4.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.0K
Evolutionary Relationships through Genome Comparisons
5.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.8K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K


