概括:一种基于潜在空间的蛋白质序列家族的生成模型
Hoda Akl1, Brooke Emison2, Xiaochuan Zhao1
1Department of Physics, University of Florida, Gainesville, Florida, United States of America.
PLoS computational biology
|November 27, 2023
概括
我们开发了GENERALIST,这是蛋白质序列的新型生成模型. 它准确地模拟蛋白质家族,即使数据有限,并产生现实的序列组合.
科学领域:
- 计算生物学是一种计算生物学.
- 蛋白质工程是一种蛋白质工程.
- 生物信息学是一种生物信息学.
背景情况:
- 生成模型对蛋白质科学至关重要,但与大型蛋白质和低覆盖范围的家族斗争.
- 现有的方法在推断,准确性和过拟合方面面临挑战.
研究的目的:
- 介绍 GENERALIST,一种新的蛋白质序列生成模型.
- 为了解决当前蛋白质建模中的生成方法的局限性.
主要方法:
- 一般主义者采用非线性张量因子化方法.
- 该模型的设计是为了方便学习,调整性和准确性.
主要成果:
- GENERALIST准确地捕获了高级氨基酸共变统计数据.
- 它预测稳定的蛋白质结构,并产生密切匹配自然的序列组合.
- 该模型为蛋白质序列创建了一个信息潜伏空间.
结论:
- GENERALIST提供了一种准确和有效的方法来建模蛋白序列变异性.
- 它克服了现有的生成模型的主要局限性.
- 这种工具将推动对蛋白质序列多样性和功能的研究.
相关概念视频
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K
Gene Families
8.8K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.8K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains
3.1K
3.1K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K


