Related Experiment Video
Updated: Feb 16, 2026

10:02
Identification of Protein Interacting Partners Using Tandem Affinity Purification
Published on: February 25, 2012
38.3K
Fast protein classification by using the most significant pairs.
1Computer Science Department, Faculty of Science and Information Technology, Zarka Private University, Zarka, Jordan.
EXCLI Journal
|December 20, 2017
Summary
This study presents a novel method for faster protein classification by rewriting sequences using significant pairs. Using 300 pairs maintains classification accuracy while speeding up the process.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Protein classification is crucial for understanding protein function and evolution.
- Current methods, like profile hidden Markov models (HMMs), can be computationally intensive.
- Accelerating protein classification is essential for large-scale genomic and proteomic analyses.
Purpose of the Study:
- To develop and evaluate a new approach for accelerating protein classification.
- To investigate the impact of sequence rewriting using significant pairs on classification speed and accuracy.
- To determine the optimal number of significant pairs for efficient and accurate protein classification.
Main Methods:
- Rewriting protein sequences by identifying and utilizing the most significant pairs.
- Analyzing the reduction in sequence length based on the number of significant pairs used (100, 200, 300).
- Comparing the classification time and accuracy of the new method against traditional profile HMMs.
Main Results:
- Sequence length reduction achieved: 0.86 (100 pairs), 0.91 (200 pairs), and 0.95 (300 pairs).
- Average time reduction observed: 0.53% (100 pairs), 0.33% (200 pairs), and 0.22% (300 pairs).
- Using at least 300 significant pairs is required to match the classification rate of profile HMMs.
Conclusions:
- The proposed method effectively speeds up protein classification testing time.
- Significant pairs can be used to reduce sequence length and computational load.
- A minimum of 300 significant pairs is recommended to maintain classification accuracy comparable to profile HMMs.
Related Concept Videos
Protein Families
17.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
17.2K
Protein Families
4.5K
4.5K
Protein Networks
4.6K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.6K
Protein-protein Interfaces
14.8K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.8K
Protein-Protein Interfaces
4.5K
4.5K
Conservation of Protein Domains Over Different Proteins
14.7K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.7K

