DeepMPSF:一个深度学习网络,基于多个蛋白质序列特征预测一般蛋白质酸化地点
Jingxin Xie1, Lijun Quan1,2,3, Xuejiao Wang1
1School of Computer Science and Technology, Soochow University, Suzhou 215006, China.
Journal of chemical information and modeling
|November 6, 2023
概括
DeepMPSF是一种新的酸化位预测模型,通过整合多个蛋白质序列特征来提高准确性. 这种先进的方法显示了跨物种的卓越性能,为生物研究提供了宝贵的见解.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
背景情况:
- 翻译后的修改,特别是酸化,对于细胞过程和疾病至关重要.
- 现有的用于预测酸化位点的计算方法往往缺乏全面的上下文信息,因为它们依赖于有限的序列特征.
研究的目的:
- 开发一个先进的计算模型,DeepMPSF,用于准确预测蛋白质酸化部位.
- 通过结合各种蛋白质序列特征来克服现有方法的局限性.
主要方法:
- DeepMPSF使用两个子网络 (S71SFE和BBFE) 来提取序列语义和蛋白质背景生物物理特征.
- 该模型采用集体学习来解决训练和预测期间不平衡的数据集.
主要成果:
- 与人类蛋白质数据集的基准方法相比,DeepMPSF在S/T和Y残留的预测性能优越.
- 该模型表现出优异的跨物种概括性,在*Mus musculus*和*Rattus norvegicus*测试集上显著改善了AUC,F1得分和MCC指标.
结论:
- DeepMPSF的多功能方法显著提高了酸化地点预测的准确性.
- 该模型为未来的酸化位点分析和下游应用研究提供了宝贵的见解.
相关概念视频
Protein Kinases and Phosphatases
13.2K
Proteins undergo chemical modifications that trigger changes in the charge, structure, and conformation of the proteins. Phosphorylation, acetylation, glycosylation, nitrosylation, ubiquitination, lipidation, methylation, and proteolysis are various protein modifications that regulate protein activity. Such modifications are usually enzyme-driven.
Protein kinases
Many proteins in the cell are regulated by phosphorylation, the addition of a phosphate group. A family of enzymes called kinases...
Protein kinases
Many proteins in the cell are regulated by phosphorylation, the addition of a phosphate group. A family of enzymes called kinases...
13.2K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Phosphorylation
50.4K
The addition or removal of phosphate groups from proteins is the most common chemical modification that regulates cellular processes. These modifications can affect the structure, activity, stability, and localization of proteins within cells as well as their interactions with other proteins.
During phosphorylation, protein kinases transfer the terminal phosphate group of ATP to specific amino acid side chains of substrate proteins. Serine, threonine, and tyrosine are the most commonly...
During phosphorylation, protein kinases transfer the terminal phosphate group of ATP to specific amino acid side chains of substrate proteins. Serine, threonine, and tyrosine are the most commonly...
50.4K
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K


