使用基于对齐和无对齐的方法预测植物抗性蛋白质
Pushpendra Singh Gahlot1, Shubham Choudhury1, Nisha Bajiya1
1Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India.
Proteomics
|November 24, 2024
概括
预测植物抗病 (PDR) 蛋白质对于作物保护至关重要. 结合机器学习和BLAST的新混合模型实现了0.98 AUROC,显著改善了PDR蛋白质的预测和设计.
科学领域:
- 植物生物学 植物生物学
- 生物信息学是一种生物信息学.
- 计算生物学是一种计算生物学.
背景情况:
- 植物抗病 (PDR) 蛋白质对于识别病原体和保护作物至关重要.
- 准确预测PDR蛋白质对于理解植物病原体相互作用和制定有效的作物保护策略至关重要.
研究的目的:
- 开发和验证一个强大的混合模型来预测和设计植物疾病耐药性蛋白质.
- 克服PDR蛋白质识别中传统的基于对齐的方法的局限性.
主要方法:
- 探索了基于对齐的方法 (BLAST,MERCI) 并发现它们不足.
- 开发了使用蛋白质组成特征的无对齐机器学习 (ML) 模型.
- 包含进化信息来提高ML模型的性能.
- 创建了一个混合/整体模型,将最好的ML模型与BLAST结合起来.
主要成果:
- 使用组合特征的ML模型实现了0.91.9的AUROC.
- 整合进化信息使ML模型的AUROC提高到0.95.
- 最终的混合模型在验证数据集上获得了高的AUROC 0.98.
- 检查的验证数据集蛋白质与训练数据相似度低于40%,以便进行可靠的评估.
结论:
- 开发的混合模型显著提高了植物疾病耐药性蛋白质的预测准确性.
- 这项研究为科学界提供了一个有价值的工具,PlantDRPpred,用于预测和设计PDR蛋白.
- 这项工作促进了植物病理学和作物改进战略的进步.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K


