EPI-SF: essential protein identification in protein interaction networks using sequence features.
Sovan Saha1, Piyali Chatterjee2, Subhadip Basu3
1Department of Computer Science & Engineering (Artificial Intelligence & Machine Learning), Techno Main Salt Lake, Kolkata, West Bengal, India.
Peerj
|March 18, 2024
Summary
This study introduces EPI-SF, a machine learning method to identify essential proteins using protein sequence features. The Random Forest model demonstrated superior performance in predicting essential proteins in yeast and human networks, outperforming existing methods.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning Applications in Proteomics
Background:
- Identifying essential proteins is crucial for understanding organism viability and function.
- Traditional methods for protein identification are time-consuming, labor-intensive, and costly.
- Computational approaches, particularly machine learning, offer a more efficient alternative.
Purpose of the Study:
- To propose and evaluate a novel machine learning-based technique, EPI-SF, for identifying essential proteins.
- To compare the performance of various machine learning models using protein sequence features.
- To apply the validated method to predict essential proteins in human protein-protein interaction networks (PPINs) and assess their potential involvement in diseases like COVID-19.
Main Methods:
- Extraction of protein sequence features from protein-protein interaction networks (PPINs).
- Application of conventional machine learning (ML) models including XGBoost, AdaBoost, Logistic Regression, SVM, Decision Tree, Random Forest, and Naïve Bayes.
- Evaluation of model performance using metrics such as precision, recall, F1-score, and AUC, with a focus on the Random Forest model.
Main Results:
- The Random Forest model achieved the highest performance in identifying essential proteins in yeast PPINs, with precision, recall, F1-score, and AUC values of 0.703, 0.720, 0.711, and 0.745, respectively.
- EPI-SF, utilizing the Random Forest model, outperformed traditional centrality measures (e.g., betweenness centrality, closeness centrality) and deep learning methods (e.g., DeepEP).
- The method was successfully applied to predict novel essential proteins in the human PPIN.
Conclusions:
- The proposed EPI-SF technique, powered by machine learning and protein sequence features, is an effective approach for identifying essential proteins.
- The Random Forest model demonstrates significant potential for essential protein prediction compared to existing methods.
- The findings suggest potential roles for identified essential proteins in disease transmission, including COVID-19, warranting further investigation.
Related Concept Videos
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Organization
6.5K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.5K
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Protein-Protein Interfaces
3.8K
3.8K


