在GST超级家族分类上测试基于嵌入的对齐的能力:蛋白质长度的作用
Gabriele Vazzana1, Castrense Savojardo1, Pier Luigi Martelli1
1Biocomputing Group, Department of Pharmacy and Biotechnology, University of Bologna, 40126 Bologna, Italy.
Molecules (Basel, Switzerland)
|October 16, 2024
概括
蛋白质语言模型准确地分类了Glutathione S-transferases (GSTs),揭示了46%的UniProt注释的GSTs缺乏正规的结构长度,影响了它们的家族分类. 该方法还对以前未分配的GST进行了分类.
科学领域:
- 生物化学 生物化学
- 生物信息学是一种生物信息学.
- 结构生物学 结构生物学
背景情况:
- 氨酸S转移酶 (GSTs) 对于细胞解毒至关重要,但由于不同的类,分类,位置,折叠和序列身份,它们存在注释挑战.
- 传统的分类依赖于序列相似性,往往未能说明远程同类物或缺乏已知的结构的蛋白质.
研究的目的:
- 评估基于蛋白质语言模型的对齐,用于分类Glutathione S-transferases (GSTs).
- 将这种新的方法与UniProt的ARBA/UNI基于规则的注释进行比较.
- 在现有的GST分类中识别潜在的不准确性.
主要方法:
- 使用基于嵌入的对齐方法,使用Meta ESM2-15b蛋白语言模型.
- 由UniProt-ARBA/UNI规则先前注释的 15,061 种 GST 蛋白质被分类.
- 分析了结构模板长度的保存,以确保分类的准确性.
主要成果:
- 基于嵌入的对齐实现了超过99%的匹配精度与UniProt的自动程序.
- 46%的UniProt自动分类的GST蛋白质偏离了规范长度,从而损害了结构模板的保存.
- 在64,207个未分配的GST蛋白中,有41%的蛋白质根据结构模板长度成功进行了分类.
结论:
- 基于蛋白质语言模型的对齐为GST分类提供了一个强大的方法,超越了传统序列相似性的局限性.
- 在现有的UniProt GST注释中发现了显著的结构性不一致性,突出了改进分类策略的必要性.
- 开发的方法有效地对以前未被分配的GST蛋白进行了分类,提高了蛋白质家族数据库的完整性.
相关概念视频
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Ligand Binding and Linkage
4.8K
Allosteric proteins have more than one ligand binding site; the binding of a ligand to any of these sites influences the binding of ligands to the other sites. When a protein is allosteric, its binding sites are called coupled or linked. In the case of enzymes, the site that binds to the substrate is known as the active site and the other site is known as the regulatory site. When a ligand binds to the regulatory site, this leads to conformational changes in the protein that can influence...
4.8K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Gene Evolution - Fast or Slow?
7.0K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.0K


