微调蛋白质嵌入用于功能相似性评估
Andrew Dickson1, Mohammad R K Mofrad1
1Departments of Bioengineering and Mechanical Engineering, Molecular Cell Biomechanics Laboratory, University of California, Berkeley, CA 94720, United States.
Bioinformatics (Oxford, England)
|July 10, 2024
概括
微调蛋白质语言模型可以改善它们的嵌入功能注释和蛋白质家族发现. 这些增强的嵌入式保持可解释性,并在关键生物信息学任务中优于标准方法.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 蛋白质科学 蛋白质科学
背景情况:
- 具有未知功能的蛋白质通常与已知的蛋白质进行比较,使用序列相似性或学习嵌入空间.
- 蛋白质序列嵌入有助于注释,聚类和发现蛋白质家族.
- 蓄意设计嵌入来增强下游任务的潜力仍未得到充分探索.
研究的目的:
- 调查微调蛋白质语言模型是否可以提高下游任务中蛋白质嵌入的质量和实用性.
- 评估微调对功能注释和蛋白质家族发现的影响.
主要方法:
- 微调预训练的蛋白质语言模型,使用基因本体学 (GO) 术语的分类损失.
- 评估微调嵌入在GO注释的K-最近邻居分类器中的性能.
- 使用聚类评估细调嵌入的质量,以发现蛋白质家族.
主要成果:
- 语言模型的直接微调显著提高了功能注释的蛋白质嵌入质量.
- 与标准方法相比,微调嵌入在GO注释中取得了更高的性能,甚至超过了直接微调的分类器.
- 改进的嵌入式通过蛋白质相似性比较保持可解释性,并在为蛋白质家族重新发现的聚类任务中表现良好.
结论:
- 微调蛋白质语言模型是改善蛋白质嵌入质量的有效策略.
- 增强的嵌入方便更准确的功能注释和蛋白质家族发现.
- 这种方法为生物信息学研究提供了一个强大的工具,它将可解释性与性能改进相结合.
相关概念视频
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Improving Translational Accuracy
9.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.9K
Protein Folding Quality Check in the RER
3.7K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.7K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Complexes with Interchangeable Parts
2.5K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.5K


