Related Experiment Video
Updated: Feb 8, 2026

06:50
Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
2.6K
Improving Prediction of Self-interacting Proteins Using Stacked Sparse Auto-Encoder with PSSM profiles.
Yan-Bin Wang1,2, Zhu-Hong You2, Li-Ping Li2
1University of Chinese Academy of Sciences, Beijing 100049, China.
International Journal of Biological Sciences
|July 11, 2018
Summary
This study introduces a new machine learning approach for identifying self-interacting proteins (SIPs), crucial for cellular processes. The method combines Zernike Moments and deep learning for efficient and accurate SIP detection.
Area of Science:
- Computational biology
- Bioinformatics
- Machine learning in protein science
Background:
- Self-interacting proteins (SIPs) are vital for cellular functions including signal transduction and gene regulation.
- Experimental methods for SIP identification are costly and time-consuming, necessitating efficient computational approaches.
Purpose of the Study:
- To develop and validate a novel computational method for accurate identification of self-interacting proteins (SIPs).
- To leverage machine learning techniques for efficient feature extraction and classification of SIPs.
Main Methods:
- Utilized Zernike Moments (ZMs) for feature extraction from Position Specific Scoring Matrix (PSSM) data.
- Employed Stacked Sparse Auto-Encoder (SSAE) for deep feature dimension reduction and noise removal.
- Applied Probabilistic Classification Vector Machines (PCVM) for the final classification of SIPs.
Main Results:
- The proposed method achieved high prediction accuracies of 92.55% on *S. cerevisiae* and 97.47% on Human SIP datasets.
- Comparative analysis demonstrated that the PCVM classifier outperformed Support Vector Machine (SVM) and other existing methods.
- The integrated approach shows significant promise for accurate and efficient SIP identification.
Conclusions:
- The novel computational strategy integrating ZMs, SSAE, and PCVM offers a powerful tool for identifying self-interacting proteins.
- This machine learning-based approach provides a cost-effective and time-efficient alternative to experimental methods for SIP detection.
- The study highlights the potential of advanced computational techniques in advancing our understanding of protein interactions and cellular mechanisms.
Related Concept Videos
Encoding
867
Information enters the brain through encoding, which is the input of information into the memory system. Once sensory information is received from the environment, the brain labels or codes it. The information is then organized with similar information and connected to existing concepts. Encoding occurs through automatic processing and effortful processing.
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
867
Predicting Molecular Geometry
46.0K
VSEPR Theory for Determination of Electron Pair Geometries
46.0K
Protein and Protein Structure
88.2K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
88.2K
Improving Translational Accuracy
15.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.0K
Protein Networks
4.6K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.6K
Prediction Intervals
3.4K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.4K

