Prediction of carbohydrate-binding proteins from sequences using support vector machines.
Seizi Someya1, Masanori Kakuta, Mizuki Morita
1Department of Biotechnology, The University of Tokyo, 1-1-1 Yayoi, Bunkyo-ku, Tokyo 113-8657, Japan.
Advances in Bioinformatics
|October 12, 2010
Summary
We developed a machine learning method to predict carbohydrate-binding proteins from amino acid sequences. This approach accurately identifies proteins interacting with sugar chains, aiding in understanding their physiological roles.
Area of Science:
- Bioinformatics
- Computational Biology
- Protein Science
Background:
- Carbohydrate-binding proteins play crucial roles in numerous physiological processes.
- Accurate identification of these proteins is essential for biological research.
Purpose of the Study:
- To develop a computational method for predicting carbohydrate-binding proteins based on amino acid sequences.
- To enhance the accuracy of identifying proteins that interact with sugar chains.
Main Methods:
- Utilized Support Vector Machines (SVMs) for sequence-based prediction.
- Constructed positive and negative datasets for training and validation.
- Evaluated prediction performance using the area under the ROC curve (0.92).
- Investigated amino acid grouping methods to optimize sequence pattern learning.
Main Results:
- Achieved a high prediction accuracy of 0.92 (Area Under ROC Curve).
- Demonstrated improved true positive rates when combining the SVM method with homology-based predictions on the H-invDB human genome database.
Conclusions:
- The developed SVM-based method provides an effective means for predicting carbohydrate-binding proteins.
- Combining computational prediction with homology-based approaches enhances identification accuracy within genomic datasets.
Related Concept Videos
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein Networks
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Protein-protein Interfaces
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a polypeptide...

