HiFun:通过一种新型的蛋白质语言自我注意力模型进行同质独立蛋白质功能的预测
Jun Wu1, Haipeng Qing1, Jian Ouyang1
1Center for Bioinformatics and Computational Biology, the Institute of Biomedical Sciences and The School of Life Sciences, East China Normal University, Shanghai , 200241, China.
我们开发了HiFun,这是一种深度学习方法,用于仅使用氨基酸序列来预测蛋白质功能. 这种方法有效地注释了缺乏对已知的序列同质性的新型蛋白质,推进了元基因组研究.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 从序列中预测蛋白质功能是具有挑战性的,特别是对于具有低同源性的元基因组学中的新型蛋白质.
- 现有的基于同质的方法与未表征的蛋白质作斗争.
研究的目的:
- 开发一种独立于序列同源性的蛋白质功能预测的新方法.
- 改进在元基因组学研究中发现的未知蛋白质的注释.
主要方法:
- 拟议的同质独立蛋白质功能注释 (HiFun) 方法使用统一的深度学习模型.
- 处理蛋白质序列作为一种语言形式来提取潜在特征.
- 使用CAFA3挑战基准数据集和指标评估HiFun的稳定性.
主要成果:
- 在UHGP-50目录中,HiFun成功注释了2,212,663种未知的蛋白质,并确定了新的图案.
- 证明了HiFun从序列中提取功能相关结构特征的能力.
- 对缺乏对已知蛋白质同质性的蛋白质实现了准确的功能注释.
结论:
- HiFun提供了一个强大的解决方案,用于注释非同源蛋白质,这对于元基因组和元转录组研究至关重要.
- 该方法增强了对各种生态领域微生物适应的理解.
- 在http://www.unimd.org/HiFun上提供了一个免费的,可访问的HiFun网络服务,以便实际使用.
更多相关视频
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
10:21Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
相关概念视频
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein-protein Interfaces
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein Families
