GPSFun:用语言模型进行几何意识的蛋白质序列功能预测
Qianmu Yuan1, Chong Tian1, Yidong Song1
1School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, Guangdong 510000, China.
Nucleic acids research
|May 13, 2024
概括
GPSFun是一个新的Web服务器,它使用语言模型和几何深度学习来进行蛋白质功能注释. 它只从序列中准确预测蛋白质功能,推动药物发现和疾病机制研究.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 结构生物学是结构生物学.
背景情况:
- 了解蛋白质功能对于疾病研究和药物向鉴定至关重要.
- 在蛋白质序列数据的快速增长和功能注释的可用性之间存在很大的差距.
- 之前的方法如GraphPPIS,GraphSite,LMetalSite和SPROF-GO已被开发用于蛋白质功能预测.
研究的目的:
- 为了介绍GPSFun,一个多功能Web服务器,用于几何意识的蛋白质序列函数注释.
- 整合语言模型和几何深度学习以增强蛋白质功能预测.
- 为各种下游功能预测提供一个用户友好的平台.
主要方法:
- 使用大型语言模型来预测3D蛋白质构造和提取序列嵌入.
- 使用几何图形神经网络来分析蛋白质图中的序列和结构模式.
- 开发一个具有直观接口和可访问性可视化的Web服务器.
主要成果:
- 与最先进的方法相比,GPSFun在各种预测任务上表现出卓越的性能.
- 服务器准确地预测了蛋白质连接体结合点,基因本体学,亚细胞位置和蛋白质溶解度.
- 有效的函数注释可以实现,而不依赖于多个序列对齐或实验结构.
结论:
- GPSFun提供了一种强大而通用的工具,用于使用基于序列的预测进行蛋白质功能注释.
- 集成先进的人工智能技术解决了在大型蛋白质序列数据集中功能注释有限的挑战.
- GPSFun是免费访问的,支持分子生物学,药物发现和疾病机制阐明的更广泛的研究.
更多相关视频
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
1.8K
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
7.3K
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Conservation of Protein Domains
3.1K
3.1K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
