EPI-SF:使用序列特征在蛋白相互作用网络中识别基本蛋白质
Sovan Saha1, Piyali Chatterjee2, Subhadip Basu3
1Department of Computer Science & Engineering (Artificial Intelligence & Machine Learning), Techno Main Salt Lake, Kolkata, West Bengal, India.
PeerJ
|March 18, 2024
概括
本研究介绍了EPI-SF,这是一种机器学习方法,使用蛋白质序列特征识别基本蛋白质. 随机森林模型在预测酵母和人体网络中的必需蛋白质方面表现出卓越的表现,优于现有方法.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 机器学习在蛋白质组学中的应用
背景情况:
- 识别必要的蛋白质对于理解生物体的生存能力和功能至关重要.
- 传统的蛋白质识别方法是耗时的,劳动密集的和昂贵的.
- 计算方法,特别是机器学习,提供了一个更有效的替代方案.
研究的目的:
- 提出和评估一种基于机器学习的新技术EPI-SF,用于识别基本蛋白质.
- 使用蛋白质序列特征比较各种机器学习模型的性能.
- 应用经过验证的方法来预测人类蛋白质-蛋白质相互作用网络 (PPIN) 中的基本蛋白质,并评估它们在COVID-19等疾病中的潜在参与.
主要方法:
- 从蛋白质-蛋白质相互作用网络 (PPIN) 中提取蛋白质序列特征.
- 应用传统的机器学习 (ML) 模型,包括XGBoost,AdaBoost,后勤回归,SVM,决策树,随机森林和naive Bayes.
- 使用精度,回忆,F1得分和AUC等指标评估模型性能,重点是随机森林模型.
主要成果:
- 随机森林模型在确定酵母PPIN中的必需蛋白质方面取得了最高的性能,精度,回忆,F1得分和AUC值分别为0.703,0.720,0.711和0.745.
- 使用随机森林模型的EPI-SF的表现优于传统的中心性措施 (例如,中间中心性,接近中心性) 和深度学习方法 (例如,DeepEP).
- 该方法成功地用于预测人类PPIN中的新型基本蛋白质.
结论:
- 拟议的EPI-SF技术,由机器学习和蛋白质序列特征提供动力,是识别必要蛋白质的有效方法.
- 与现有方法相比,随机森林模型显示了基本蛋白质预测的巨大潜力.
- 这些发现表明,在疾病传播 (包括COVID-19) 中,已识别的必需蛋白质可能发挥作用,因此需要进一步调查.
相关概念视频
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Organization
6.5K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.5K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Protein-Protein Interfaces
3.8K
3.8K


