使用进化规模模型在蛋白质中O-GlcNAc修饰的特定地点预测
Ayesha Khalid1, Afshan Kaleem1, Wajahat Qazi2
1Department of Biotechnology, Lahore College for Women University, Lahore, Pakistan.
PloS one
|December 31, 2024
概括
这项研究引入了进化规模模型2 (ESM-2),用于预测人类蛋白质中的O-GlcNAc位点. 通过避免过度匹配和加强糖蛋白质组研究,ESM-2显示出有效的预测,优于传统模型.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 蛋白质组学是指蛋白质组学.
背景情况:
- 蛋白质糖化,特别是O-GlcNAcylation,是影响生物过程和疾病的关键翻译后修饰.
- 计算方法,包括机器学习和蛋白质语言模型,越来越多地用于预测O-GlcNAc站点,提供效率和降低成本.
研究的目的:
- 评估进化尺度模型2 (ESM-2) 在预测人类蛋白质中O-GlcNAc糖化位点方面的有效性.
- 使用ESM-2建立O-GlcNAc站点预测的计算方法,解决现有文献中的差距.
主要方法:
- 使用来自O-GlcNAc数据库的约1100个O-链接糖蛋白序列进行模型训练.
- 采用ESM-2模型,一种蛋白质语言模型,用于预测人类蛋白质中的O-GlcNAc位点.
- 将ESM-2的表现与传统模型进行比较,以评估准确性和过度拟合趋势.
主要成果:
- 在培训期间,ESM-2模型显示了持续的改进,准确率为78.30%,回忆率为78.30%,精度为61.31%,F1得分为68.74%.
- 通过表现出最佳的培训和测试预测,并避免显著的过度装配,ESM-2表现出比传统模型更好的性能.
- 该模型在预测人类蛋白质中的O-GlcNAc位点方面的有效性得到了验证.
结论:
- 在人类蛋白质中,ESM-2模型对于准确的O-GlcNAc位点预测是有效的.
- 准确的O-GlcNAc位点预测可以显著推进糖蛋白质组研究,有助于了解蛋白质功能,疾病机制和治疗发展.
- 未来的研究应该探索多样化的数据,更长的序列和增强的计算资源,以进一步完善预测模型.
更多相关视频
05:57Author Spotlight: In Silico Creation and Impact of Carbonylated Amino Acids on Protein Structure and Function
Published on: April 26, 2024
286
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
7.2K
相关概念视频
Conserved Binding Sites
4.1K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.1K
Conservation of Protein Domains Over Different Proteins
10.5K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.5K
Protein Glycosylation
6.4K
Glycosylation, the most common post-translational modification for proteins, serves diverse functions. Adding sugars to proteins makes the proteins more resistant to proteolytic digestion. Glycosylated proteins can act as markers and receptors to promote cell-cell adhesion. Additionally, they have many essential quality control functions in the cell, such as correct protein folding and facilitating transport of misfolded proteins to the cytosol, which can be degraded.
Glycosylation occurs in...
Glycosylation occurs in...
6.4K
Protein Modifications in the RER
4.8K
Modification of secretory and transmembrane proteins entering the rough ER begins in the ER lumen. These modifications aid in protein folding and stabilize the acquired tertiary structure. Protein modifications in the rough ER co-occur at different stages of protein folding.
Broadly, these modifications can be categorized into four main categories — glycosylation, formation of disulfide bonds, assembly of protein subunits, and specific proteolytic cleavages like removal of signal...
Broadly, these modifications can be categorized into four main categories — glycosylation, formation of disulfide bonds, assembly of protein subunits, and specific proteolytic cleavages like removal of signal...
4.8K
Oligosaccharide Assembly
2.7K
Protein glycosylation starts in the ER lumen and continues in the Golgi apparatus. Glycosyltransferases catalyze the addition of sugar molecules or glycosylation of proteins. Usually, these enzymes add sugars to the hydroxyl groups of selected serine or threonine residues to form O-linked glycans or the amino groups of asparagine residues to form N-linked glycans. Different positions on the same polypeptide chain can contain differently linked glycans.
Multiple sugar molecules that may or may...
Multiple sugar molecules that may or may...
2.7K
Protein Folding Quality Check in the RER
3.6K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.6K
