CoNglyPred:使用ESM-2和结构特征与图形网络和协同注意力准确预测N链接的糖化位点
Hongmei Wang1,2, Long Zhao1,2, Ziyuan Yu1,2
1Department of Mathematics, School of Mathematics and Computer Sciences, Nanchang University, Nanchang, China.
Proteomics
|October 3, 2024
概括
CoNglyPred通过整合蛋白质序列和3D结构信息,准确地预测N链接的糖化位点. 这种新的计算方法提高了糖化分析的效率和精度.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 蛋白质组学是指蛋白质组学.
背景情况:
- 与N结合的糖化对蛋白质功能至关重要,影响折叠,免疫力和运输.
- 实验性地确定糖基化位点是费力和耗时的.
- 现有的计算方法很难有效地整合序列和3D结构数据.
研究的目的:
- 开发一个高精度的计算模型,用于预测N链接的糖化位.
- 利用最近在蛋白质语言模型和结构预测方面的进展.
- 改进序列和结构信息的整合,以提高预测.
主要方法:
- 使用ESM-2蛋白语言模型进行序列嵌入.
- 采用图形变压器网络来处理来自AlphaFold2.2.的3D蛋白质结构.
- 综合序列和结构信息,使用共同注意力机制.
主要成果:
- 与独立测试数据集上的最先进模型相比,ConglyPred在独立测试数据集上表现出更高的性能.
- 该模型在案例研究中表现出了卓越的表现.
- 介绍了第一个关于N-链接甘氨基化预测因子的不确定性量化报告.
结论:
- CoNglyPred在预测N链接的糖化位点方面取得了重大进展.
- 该模型有效地整合了序列和结构数据,克服了以前方法的局限性.
- 不确定性量化为预测可靠性提供了宝贵的见解.
相关概念视频
Oligosaccharide Assembly
3.8K
Protein glycosylation starts in the ER lumen and continues in the Golgi apparatus. Glycosyltransferases catalyze the addition of sugar molecules or glycosylation of proteins. Usually, these enzymes add sugars to the hydroxyl groups of selected serine or threonine residues to form O-linked glycans or the amino groups of asparagine residues to form N-linked glycans. Different positions on the same polypeptide chain can contain differently linked glycans.
Multiple sugar molecules that may or may...
Multiple sugar molecules that may or may...
3.8K
Protein Glycosylation
10.4K
Glycosylation, the most common post-translational modification for proteins, serves diverse functions. Adding sugars to proteins makes the proteins more resistant to proteolytic digestion. Glycosylated proteins can act as markers and receptors to promote cell-cell adhesion. Additionally, they have many essential quality control functions in the cell, such as correct protein folding and facilitating transport of misfolded proteins to the cytosol, which can be degraded.
Glycosylation occurs in...
Glycosylation occurs in...
10.4K
Conserved Binding Sites
5.3K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.3K
Proteoglycans
5.2K
Glycans, a class of complex heterogeneous molecules, can be covalently attached to proteins to form glycosylated proteins that regulate various physiological and pathological processes. Glycosylated proteins or glycoproteins comprise N-linked and O-linked oligosaccharides. O-glycosylation is the most common type of protein glycosylation. Here, glycans attach to the oxygen atom of the hydroxyl groups of Serine or Threonine residues. O-linked glycosylation occurs later in protein processing,...
5.2K
Protein Folding Quality Check in the RER
5.6K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
5.6K


