PSSM-Sumo:基于深度学习的智能模型,用于使用歧视性特征预测聚合点
Salman Khan1, Salman A AlQahtani2, Sumaiya Noor3
1Department of Computer Science, Abdul Wali Khan University Mardan, Mardan, KPK, Pakistan.
BMC bioinformatics
|August 30, 2024
概括
这项研究介绍了PSSM-Sumo,这是一种新型的深度学习模型,用于准确预测化位点,对于了解蛋白质功能和帕金森氏症和阿尔茨海默氏症等疾病至关重要. 该模型实现了98.71%的准确性,超过了现有方法.
科学领域:
- 蛋白质组学和计算生物学
- 分子生物学和生物化学 分子生物学和生物化学
背景情况:
- 翻译后修饰 (PTMs),包括sumoylation,调节关键的细胞过程和蛋白质功能.
- 苏莫化位点的识别对于了解蛋白质组功能和帕金森病和阿尔茨海默病等疾病至关重要.
- 现有的计算模型用于预测化部位,但在传统的学习方法上存在局限性.
研究的目的:
- 开发一个强大的计算模型,用于准确预测化部位.
- 通过优化深度学习和特征提取,提高识别化部位的有效性.
- 增强对化在疾病中的作用的理解,并促进药物发现.
主要方法:
- 引入伪位置特定评分矩阵 (PsePSSM) 模型,PSSM-Sumo.
- 使用支向量机 (SFS-SVM) 实现顺序向前选择,以实现最佳的特征选择.
- 使用多层深度神经网络 (DNN) 作为核心分类器.
- 使用十倍交叉验证和统计指标 (MCC,精度,灵敏度,特异性,AUC) 的性能评估.
主要成果:
- PSSM-Sumo表现出极高的预测准确度,平均达到98.71%的预测准确度.
- 该模型显著超过了现有的化部位预测方法的性能.
- 使用SFS-SVM优化了功能选择,简化了计算过程并删除了无关的功能.
结论:
- PSSM-Sumo模型提供了一个非常准确和强大的方法来识别化部位.
- 这一进步有望加速药物发现并改善与化相关疾病的诊断.
- 这项研究强调了深度学习在计算生物学中的潜力,用于预测关键蛋白质修饰.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Cis-regulatory Sequences
3.0K
3.0K
lncRNA - Long Non-coding RNAs
2.8K
2.8K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Peptide Identification Using Tandem Mass Spectrometry
6.4K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.4K


