VF-Fuse:一种双路功能融合和代更新架构,用于病毒性因素预测
Liang Huang1,2, Xiangyu Yu1,2, Shumei Li1,2
1School of Information and Artificial Intelligence, Anhui Agricultural University, Changjiang West Road 130, Shushan District, Hefei 230036, Anhui Province, China.
Briefings in bioinformatics
|September 22, 2025
概括
使用新型深度学习融合方法改进了对细菌毒性因子 (VF) 的预测. 我们的VF-Fuse模型提高了更好的传染病控制的准确性.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 在基因组学中的机器学习.
背景情况:
- 对细菌毒性因子 (VFs) 的准确预测对于传染病管理至关重要.
- 传统的方法难以捕获复杂的蛋白质序列属性.
- 大规模蛋白质语言模型 (PLM) 提供高级特征表示.
研究的目的:
- 开发一种用于增强细菌毒性因子预测的新框架.
- 克服来自不同PLM的简单特征连接的局限性.
- 为VF识别创建一个强大而准确的模型.
主要方法:
- 来自ESM-2和ProtT5蛋白语言模型的工程特征.
- 开发了新的架构:用于功能增强的VF-Iter和用于智能嵌入集成的双路功能融合 (DPF).
- 通过两阶段的过程构建了VF-Fuse模型:选择基准模型和对比组合技术 (多数投票).
主要成果:
- 在一个独立的测试组件上,VF-Fuse实现了新的最先进的性能.
- 获得了87.15%的F1评分 (3.3%的改善) 和73.61%的马修斯相关系数.
- 经过DPF模型的解释性分析验证的高敏感性 (90.1%) 和特异性 (83.33%).
结论:
- 使用DPF网络和多数投票组合的VF-Fuse模型显著提高了细菌毒性因子的预测.
- 开发的融合战略有效地整合了不同PLM的互补特征.
- 这项工作通过改进的VF识别提供了对抗传染病的强大工具.
相关概念视频
Viral Recombination
24.9K
Cells are sometimes infected by more than one virus at once. When two viruses disassemble to expose their genomes for replication in the same cell, similar regions of their genomes can pair together and exchange sequences in a process called recombination. Alternatively, viruses with segmented genomes can swap segments in a process called reassortment.
24.9K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K
Viral Mutations
39.6K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
39.6K
Antigen Processing Pathways
2.1K
MHC molecules are key players in the immune response, enabling T cells to recognize and respond to specific antigens. They are present on the surface of all nucleated cells in the body and are instrumental in presenting antigens to T cells and activating them. T cells recognize the MHC-antigen complex and initiate an immune response. MHC class I and MHC class II are two main types of MHC molecules, each associated with a distinct antigen processing pathway.
MHC Class I: Presenting Endogenous...
MHC Class I: Presenting Endogenous...
2.1K
Tagging and Fusion Proteins
8.3K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
8.3K


