格莱普雷德:通过CCU-LightGBM-BiLSTM框架与多头注意力机制预测氨酸糖化位点
Yun Zuo1, Bangyi Zhang1, Yinkang Dong1
1School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi 214000, China.
Journal of chemical information and modeling
|August 9, 2024
概括
本研究介绍了Glypred,这是一种机器学习模型,可以使用先进的技术准确地预测蛋白质糖化位点,例如ClusterCentroids Undersampling (CCU) 和双向长短期记忆网络 (BiLSTM). 格莱普雷德提供了一个强大的工具来识别这些网站,帮助疾病研究.
科学领域:
- 生物信息学和计算生物学
- 后翻译修改 后翻译修改
- 机器学习在蛋白质学中的机器学习
背景情况:
- 蛋白质糖化是一种对氨酸和氨酸残留物的翻译后修饰,与阿尔茨海默氏症,糖尿病和动脉样硬化等疾病有关.
- 鉴定糖化位点的传统实验方法是劳动密集型和耗时的.
- 机器学习提供了一种精简的方法来预测蛋白质糖化位点,但面临着数据不平衡和特征冗余等挑战.
研究的目的:
- 开发一个准确和强大的计算模型来预测氨酸糖化位.
- 为了解决数据不平衡和特征冗余问题在糖化部位预测.
- 为研究人员提供一个用户友好的工具,以帮助识别糖化位点.
主要方法:
- 特性工程:为全面的蛋白质信息编码选择了五种不同的特性类型 (AAC,KMER,DR,PWAA,EBGW).
- 数据平衡:使用ClusterCentroids低采样 (CCU) 通过删除冗余的非糖化样本来解决不平衡的数据集.
- 功能选择:利用LightGBM算法识别和选择最重要的功能,减少维度和提高模型效率.
- 分类模型:实施了双向长期短期记忆网络 (BiLSTM),具有多头注意力机制,用于准确的站点分类,包括规范化和退出,以防止过度拟合.
主要成果:
- 开发的Glypred模型在预测 lysine glycation 位点方面取得了最佳性能.
- 将CCU,LightGBM和BiLSTM与注意力结合起来,显著提高了预测准确性和稳定性.
- 使用PyQt5成功开发了用于氨酸糖化部位预测的软件工具,提高了研究效率.
结论:
- 格莱普雷德在识别蛋白质糖化位点方面表现出高精度和稳定性,为生物信息学提供了有价值的计算工具.
- 该研究强调了整合数据平衡,特征选择和高级深度学习架构用于蛋白质修饰预测的有效性.
- 开发的软件和可用的数据集为研究人员提供了辅助选工具,减少了实验工作量,提高了与糖化相关研究的效率.
更多相关视频
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
1.8K
12:49Quantification of Site-specific Protein Lysine Acetylation and Succinylation Stoichiometry Using Data-independent Acquisition Mass Spectrometry
Published on: April 4, 2018
11.6K
相关概念视频
Ligand Binding and Linkage
4.8K
Allosteric proteins have more than one ligand binding site; the binding of a ligand to any of these sites influences the binding of ligands to the other sites. When a protein is allosteric, its binding sites are called coupled or linked. In the case of enzymes, the site that binds to the substrate is known as the active site and the other site is known as the regulatory site. When a ligand binds to the regulatory site, this leads to conformational changes in the protein that can influence...
4.8K
Ligand Binding Sites
12.8K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
12.8K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
