Improved general regression network for protein domain boundary prediction.
Paul D Yoo1, Abdur R Sikder, Bing Bing Zhou
1Advanced Networks Research Group, School of Information Technologies (J12), The University of Sydney, NSW 2006, Australia. dyoo4334@it.usyd.edu.au
BMC Bioinformatics
|March 20, 2008
Summary
A new machine learning model, IGRN, accurately predicts protein domain boundaries using PSSM and structural data. It offers superior performance and efficiency compared to existing methods.
Area of Science:
- Proteomics and Bioinformatics
- Computational Biology
- Structural Biology
Background:
- Protein domains are crucial for understanding protein structure and function.
- Existing protein domain boundary prediction methods often rely on complex machine learning techniques.
- Accurate domain boundary prediction is essential for deciphering protein architecture and function.
Purpose of the Study:
- To introduce a novel machine learning model, IGRN, for enhanced protein domain boundary prediction.
- To develop a computationally efficient model for reliable domain boundary classification.
- To improve the accuracy and generalizability of protein domain identification.
Main Methods:
- The IGRN model was trained using Position Specific Scoring Matrix (PSSM), secondary structure, and solvent accessibility information.
- An inter-domain linker index was incorporated into the model for detecting potential domain boundaries.
- The model utilizes a novel machine learning approach for classification.
Main Results:
- The IGRN model achieved 67% average prediction accuracy on the Benchmark_2 dataset for multi-domain proteins.
- It demonstrated superior predictive performance and generalization compared to widely used neural network models.
- On the CASP7 benchmark dataset, IGRN showed comparable performance to established predictors like DOMpro, achieving 70.10% accuracy.
Conclusions:
- The IGRN model favorably compares to existing machine learning-based methods and domain boundary predictors.
- It excels in identifying protein domain boundaries, demonstrating advantages in model bias, generalization, and computational efficiency.
- The study highlights IGRN's potential for accurate and efficient protein domain analysis.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Folding Quality Check in the RER
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Protein-Protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

