Adaptive compressive learning for prediction of protein-protein interactions from primary sequence
Ya-Nan Zhang1, Xiao-Yong Pan, Yan Huang
1Department of Automation, Shanghai Jiao Tong University, and Key Laboratory of System Control and Information Processing, Ministry of Education of China, Shanghai 200240, China.
Journal of Theoretical Biology
|June 4, 2011
Summary
This study introduces a novel compressed sensing method to predict protein-protein interactions (PPIs) using only amino acid sequences. This approach effectively handles high-dimensional data and reduces redundancy for more accurate yeast PPI prediction.
Area of Science:
- Computational Biology
- Bioinformatics
- Systems Biology
Background:
- Protein-protein interactions (PPIs) are crucial for biological processes.
- Predicting PPIs using only amino acid sequences is challenging due to high dimensionality and data redundancy.
- Existing sequence-based methods often suffer from overfitting and high computational complexity.
Purpose of the Study:
- To develop a novel computational approach for predicting protein-protein interactions (PPIs) in yeast Saccharomyces cerevisiae using primary amino acid sequences.
- To address the limitations of traditional sequence-based prediction methods, including high dimensionality and feature vector redundancy.
- To leverage compressed sensing theory for efficient and accurate PPI prediction from sequence data.
Main Methods:
- A novel computational approach based on compressed sensing theory was proposed.
- The method compresses high-dimensional protein sequential feature vectors into a lower-dimensional space, exploiting signal sparsity.
- The compressed signal can be reconstructed from fewer measurements compared to traditional sampling theories.
Main Results:
- The proposed compressed sensing method achieved promising results in predicting yeast Saccharomyces cerevisiae PPIs.
- The algorithm effectively compresses high-dimensional feature vectors, reducing redundancy and mitigating overfitting.
- Experimental results demonstrate the method's power in analyzing noisy biological data.
Conclusions:
- The developed compressed sensing method offers a powerful strategy for predicting PPIs from primary amino acid sequences.
- This approach effectively handles high-dimensional biological data and reduces feature vector redundancy.
- The method shows great potential for extension to other complex biological systems and analyses.
Related Concept Videos
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein-Protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.


