Related Experiment Video
Updated: Feb 14, 2026

09:47
Spotting Cheetahs: Identifying Individuals by Their Footprints
Published on: May 1, 2016
15.4K
O-GlcNAcPRED-II: an integrated classification algorithm for identifying O-GlcNAcylation sites based on fuzzy
Cangzhi Jia1, Yun Zuo1, Quan Zou2
1Department of Mathematics, Dalian Maritime University, Dalian, China.
Bioinformatics (Oxford, England)
|February 9, 2018
Summary
This study introduces O-GlcNAcPRED-II, an improved computational tool for identifying protein O-GlcNAcylation sites. The model enhances prediction accuracy and sensitivity, crucial for understanding diseases linked to O-GlcNAcylation.
Area of Science:
- Biochemistry and Molecular Biology
- Computational Biology
- Bioinformatics
Background:
- Protein O-GlcNAcylation (O-GlcNAc) is a critical post-translational modification impacting cellular processes.
- Aberrant O-GlcNAcylation is implicated in diseases like cancer and neurodegeneration.
- Existing computational methods for identifying O-GlcNAcylation sites lack sufficient prediction sensitivity.
Purpose of the Study:
- To develop an accurate and effective automated computational method for identifying protein O-GlcNAcylation sites.
- To improve upon the sensitivity and overall performance of existing prediction tools.
Main Methods:
- Developed an ensemble model, O-GlcNAcPRED-II, integrating K-means principal component analysis oversampling (KPCA) and fuzzy undersampling (FUS).
- Employed a rotation forest classifier integrating random forest, k-nearest neighbour, naive Bayesian, and support vector machine classifiers.
- Utilized feature space partitioning with four distinct sub-classifiers.
Main Results:
- O-GlcNAcPRED-II achieved high performance metrics: 81.05% sensitivity, 95.91% specificity, 91.43% accuracy, and a Matthew's correlation coefficient of 0.7928.
- The model demonstrated superior performance compared to five existing prediction tools on independent datasets.
- The ensemble approach effectively balanced positive and negative training samples.
Conclusions:
- O-GlcNAcPRED-II represents a significant advancement in the computational prediction of protein O-GlcNAcylation sites.
- The developed method offers improved accuracy and sensitivity, aiding research into O-GlcNAc-related diseases.
- The ensemble strategy and data balancing techniques are effective for enhancing predictive model performance.
Related Concept Videos
Conserved Binding Sites
5.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.2K
Conserved Binding Sites
2.0K
2.0K
Ligand Binding Sites
15.3K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
15.3K
Ligand Binding Sites
8.9K
8.9K
Trial and Error and Algorithm
435
A problem-solving strategy is a plan of action used to find a solution. Different strategies have distinct action plans. Trial and error involves trying different solutions until one works. For instance, to fix a broken printer, you might check ink levels, ensure the paper tray isn't jammed, and verify the printer's connection to your laptop. This method can be time-consuming but is commonly used. Thomas Edison, for example, used trial and error to find a suitable filament for the light...
435
Classification of Titrimetric Analysis Based on Reaction Types
1.9K
Titrimetric analysis in solution chemistry involves measuring the volume of solutions and is often called volumetric analysis. The standard solution of known concentration in the burette is called the titrant, whereas the solution of unknown concentration in the flask is called the analyte, or titrand. Titrimetric analyses can be classified into four types based on the reactions between the titrant and analyte.
Titrations between an acid and a base lead to neutralization reactions that form...
Titrations between an acid and a base lead to neutralization reactions that form...
1.9K

