ET-Pfam: 组合转移学习用于蛋白质家族预测
Sofia A Duarte1, Rosario Vitale1, Sofia Escudero1
1Research Institute for Signals, Systems and Computational Intelligence, sinc(i), FICH-UNL, CONICET, Ciudad Universitaria UNL, Santa Fe, 3000, Argentina.
ET-Pfam使用转移学习和深度学习 (DL) 模型组合来改善Pfam数据库中的蛋白质功能家族注释. 与单个DL模型和最先进的方法相比,这种新的方法显著降低了错误率.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 计算式蛋白质注释面临着挑战,因为快速序列生成超过了手动策划能力.
- Pfam数据库使用了隐藏的马尔科夫模型 (pHMMs) 来进行蛋白质家族注释,但个体模型训练错过了跨家族模式.
- 现有的深度学习 (DL) 模型在基本输入表示上提供了有限的改进.
研究的目的:
- 开发一种先进的计算方法,用于在Pfam数据库中预测蛋白质功能家族.
- 为了提高预测准确性,利用转移学习和整体策略.
- 为了解决当前蛋白质注释技术的局限性.
主要方法:
- 实施了一种新的方法,ET-Pfam,结合了转移学习和多个DL分类器的集合.
- 使用从蛋白质大语言模型中学习到的表示来训练基础DL模型.
- 采用集体策略的综合基准模型,包括每种模型和每种Pfam家族的新型学习重量投票方法.
主要成果:
- 与单个DL模型相比,ET-Pfam始终降低了错误率,大大提高了预测性能.
- 通过家庭投票策略学习的权重以7.00%的错误率实现了最佳表现.
- 这一结果大大超过了最好的个人基准模型的错误率 (12.91%) 和最先进的竞争对手.
结论:
- ET-Pfam为Pfam中的蛋白质功能家族预测提供了一种卓越的方法.
- 集体策略,特别是通过家庭投票学习的权重,在提高DL模型性能方面是有效的.
- 开发的方法解决了计算蛋白质注释的关键挑战.
更多相关视频
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
相关概念视频
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Families
Protein Families
Conservation of Protein Domains
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Multi-pass Transmembrane Proteins and β-barrels
α-Helix containing multi-pass transmembrane proteins
Multi-pass transmembrane proteins such as...
