SAFPred:使用蛋白质嵌入物对细菌的合成感知基因功能预测
Aysun Urhan1,2, Bianca-Maria Cosma1, Ashlee M Earl2
1Delft Bioinformatics Lab, Delft University of Technology Van Mourik, Delft XE 2628, The Netherlands.
Bioinformatics (Oxford, England)
|May 22, 2024
概括
我们开发了SAFPred,这是一种使用先进的蛋白质语言模型和细菌基因合成来预测细菌基因功能的新工具. SAFPred提高了准确性,特别是对于新型蛋白质,并在临床威胁细菌中发现了新的毒素.
科学领域:
- 基因组学和生物信息学
- 微生物基因组学 微生物基因组学
- 计算生物学 计算生物学
背景情况:
- 由于目前预测方法的局限性,大量细菌蛋白序列缺乏功能性注释.
- 现有的基因功能预测算法主要集中在真核生物上,并依赖于序列相似性,而这种相似性通常在新型细菌蛋白中不存在.
- 细菌的遗传学和代谢多样性需要专门的基因功能预测工具.
研究的目的:
- 开发一种改进的基因功能预测方法,专门为细菌基因组量身定制.
- 利用来自先进语言模型和细菌保护合成的蛋白质嵌入来提高注释准确性.
主要方法:
- 开发了SAFPred,这是一种新的合成感知基因功能预测工具,利用最先进的蛋白质语言模型中的蛋白质嵌入.
- 将保存的合成和细菌操作结构纳入预测模型.
- 评估了SAFPred的性能与传统的基于序列的方法和多种细菌物种的现有最先进的方法相比.
主要成果:
- 与传统和最先进的方法相比,SAFpred在细菌基因功能预测方面表现优越.
- 该工具在检测远方同类时取得了很高的准确性,即使序列相似性低至40%.
- 在SAFPred对Enterococcus物种的应用中,发现了11种假定的新毒素,突出了其发现临床相关基因的潜力.
结论:
- SAFpred代表了细菌基因功能预测的重大进步,超过了现有的方法.
- 蛋白质嵌入和合成信息的整合有效地解决了标注新型细菌蛋白质的挑战.
- 在临床威胁细菌中发现新型毒素强调了该工具在推进人类和动物健康研究方面的实用性.
相关概念视频
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Cytoskeletal Proteins in Bacteria
3.4K
Bacterial cells were initially considered simple, randomly organized structures lacking a cytoskeleton. However, the discovery of cytoskeleton homologs in bacteria led to the change of this opinion. Bacterial cytoskeletal filaments regulate the cell shape, cell polarity, cell division, and partitioning of plasmids during cell division. It was later discovered that bacterial cytoskeletal proteins, mainly actin and tubulin homologs, are diverse compared to their eukaryotic counterparts. On the...
3.4K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
The Central Dogma
21.6K
The central dogma explains the flow of genetic information from DNA nucleotides to the amino acid sequence of proteins.
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
21.6K


