通过将多种类型的生物学知识与蛋白质语言模型和图形神经网络融合来提高蛋白质功能预测
IEEE transactions on computational biology and bioinformatics
|August 14, 2025
概括
DeepFMB是一个新的深度学习框架,通过整合进化和蛋白质-蛋白质相互作用数据来增强蛋白质功能预测. 这种方法显著优于现有的方法,特别是在预测罕见的蛋白质功能方面.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 机器学习在生物学中的应用
背景情况:
- 准确的蛋白质功能注释对于理解生物过程和疾病机制至关重要.
- 计算方法为蛋白质功能预测提供了实验方法的替代方案.
- 现有的方法往往忽略了没有相互作用的蛋白质,或者依赖慢进化信息提取.
研究的目的:
- 开发一个新的深度学习框架,DeepFMB,用于准确的蛋白质功能预测.
- 通过结合多种类型的生物知识来解决现有方法的局限性.
- 改善蛋白质功能的预测,包括低频基因本体学 (GO) 术语.
主要方法:
- DeepFMB利用预训练的蛋白质语言模型来提取进化信息.
- 图形神经网络被用来生成来自蛋白质与蛋白质相互作用 (PPI) 和正义网络的特征.
- 多种类型的生物特征 (进化,PPI,正义) 的自适应融合,用于功能预测.
主要成果:
- 在F-max和AUPR指标方面,DeepFMB的表现优于八种最先进的方法.
- 该模型在预测蛋白质功能方面表现出卓越的准确性,特别是低频GO术语.
- 废弃性研究证实了综合多种类型的生物知识对预测准确性的高度相关性.
结论:
- DeepFMB提供了一种强大而准确的计算方法来预测蛋白质功能.
- 整合各种生物知识,包括进化和相互作用数据,对于推进预测能力至关重要.
- 该框架显示了加速生物发现和治疗开发的前景.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
578
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
2.0K
相关概念视频
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Protein-protein Interfaces
13.2K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
13.2K
Tagging and Fusion Proteins
6.9K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.9K
Ligand Binding Sites
13.2K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
13.2K
Protein Families
15.7K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.7K
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
