Related Experiment Video
Updated: Jul 11, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
596
EMNGly: predicting N-linked glycosylation sites using the language models for feature extraction
Xiaoyang Hou1,2, Yu Wang3, Dongbo Bu1,2
1Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Beijing 100190, China.
Bioinformatics (Oxford, England)
|November 6, 2023
Summary
Predicting N-linked glycosylation sites is crucial for understanding protein function and disease. A new method, EMNGly, uses advanced models to accurately identify these sites, outperforming existing techniques.
Area of Science:
- Biochemistry and Molecular Biology
- Computational Biology
- Bioinformatics
Background:
- N-linked glycosylation is a vital post-translational modification impacting protein folding, stability, and function.
- Dysregulation of N-linked glycosylation is linked to various diseases, highlighting the need for accurate site identification.
- Experimental identification of N-linked glycosylation sites is complex, driving the need for computational approaches.
Purpose of the Study:
- To develop and evaluate a novel computational approach for predicting N-linked glycosylation sites.
- To leverage pretrained protein language and structure models for enhanced feature extraction in N-linked glycosylation site prediction.
Main Methods:
- The EMNGly approach was developed, integrating features from pretrained Evolutionary Scale Modeling (ESM) and Inverse Folding Model (IFM).
- A support vector machine (SVM) classifier was employed for predicting N-linked glycosylation sites based on extracted features.
- Rigorous evaluation was performed using ten-fold cross-validation and independent test sets.
Main Results:
- EMNGly demonstrated superior performance compared to existing N-linked glycosylation site prediction methods.
- The approach achieved high predictive accuracy, with a Matthews Correlation Coefficient (MCC) of 0.8282.
- Exceptional performance metrics were recorded on an independent test set, including sensitivity (0.9343), specificity (0.8934), and accuracy (0.9143).
Conclusions:
- EMNGly represents a significant advancement in computational N-linked glycosylation site prediction.
- The integration of advanced protein language and structure models enhances prediction accuracy.
- This method provides a valuable tool for researchers studying protein modification and glycosylation-related diseases.
More Related Videos
Related Concept Videos
Oligosaccharide Assembly
2.9K
Protein glycosylation starts in the ER lumen and continues in the Golgi apparatus. Glycosyltransferases catalyze the addition of sugar molecules or glycosylation of proteins. Usually, these enzymes add sugars to the hydroxyl groups of selected serine or threonine residues to form O-linked glycans or the amino groups of asparagine residues to form N-linked glycans. Different positions on the same polypeptide chain can contain differently linked glycans.
Multiple sugar molecules that may or may...
Multiple sugar molecules that may or may...
2.9K
Protein Glycosylation
7.0K
Glycosylation, the most common post-translational modification for proteins, serves diverse functions. Adding sugars to proteins makes the proteins more resistant to proteolytic digestion. Glycosylated proteins can act as markers and receptors to promote cell-cell adhesion. Additionally, they have many essential quality control functions in the cell, such as correct protein folding and facilitating transport of misfolded proteins to the cytosol, which can be degraded.
Glycosylation occurs in...
Glycosylation occurs in...
7.0K
Ligand Binding Sites
12.9K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
12.9K

