AHAPC:多源特征融合和集体学习,用于多类极端爱好蛋白质预测
1AIEN Institute, Shanghai Ocean University, Shanghai, 201306, China.
Analytical biochemistry
|November 5, 2025
概括
我们开发了AHAPC,这是一种识别极端友蛋白的计算工具. 这一框架加速了对酸性,性和性蛋白质的发现,这些蛋白质对工业应用非常有价值.
科学领域:
- 生物化学 生物化学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 极端性蛋白质 (酸性,性,性) 对于工业应用至关重要.
- 对这些蛋白质的实验鉴定是资源密集型和耗时的.
研究的目的:
- 引入AHAPC,这是一个统一的计算框架,用于极端友蛋白的多类分类.
- 为了使极端友好型蛋白质的高效和可靠的计算发现.
主要方法:
- 构建一个新的基准数据集,TriExtrem.
- 蛋白质特征的提取和融合,包括预训练的蛋白质语言模型 (PLM) 嵌入和手工制作的描述符 (PSSM,序列特征).
- 深度学习模型的应用:CNN,GRU和BiLSTM用于分类任务.
主要成果:
- 在分类极端友好蛋白质方面,AHAPC表现出强的表现.
- 该框架提供可解释的预测,有助于发现过程.
- 使用多分支BiLSTM架构成功进行多类分类.
结论:
- AHAPC提供了一种强大而高效的计算方法来识别极端友好型蛋白质.
- 该框架有助于可靠地发现具有工业价值的极端类动物.
- 这项工作推进了极端爱好者研究的生物信息学领域.
相关概念视频
Tagging and Fusion Proteins
8.3K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
8.3K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Conservation of Protein Domains
3.9K
3.9K
Multi-species Conserved Sequences
4.6K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.6K
Protein Families
16.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.6K
Protein Families
4.2K
4.2K


