PRMxAI:基于使用可解释的人工智能的氨基酸空间分布来预测蛋白质氨酸甲基化位点
Monika Khandelwal1, Ranjeet Kumar Rout2
1Computer Science and Engineering Department, National Institute of Technology Srinagar, Hazratbal, Srinagar, J&K, 190006, India.
BMC bioinformatics
|October 4, 2023
概括
一种新的计算方法,PRMxAI,使用机器学习准确地预测氨酸甲基化位点. 这种方法为了解蛋白质功能的实验方法提供了更快,更有效的替代方案.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
背景情况:
- 蛋白质甲基化是调节细胞功能的关键翻译后修饰.
- 氨酸甲基化对于基因调节和信号转导等过程至关重要.
- 实验预测方法昂贵且耗时.
研究的目的:
- 开发一种新的计算方法,用于预测阿尔金甲基化位点.
- 建立一个高效和准确的替代实验预测技术.
主要方法:
- 开发了基于机器学习的预测器PRMxAI.
- 提取的基于序列的特征:双组成,物理化学性质,氨基酸组成和基于信息理论的特征.
- 利用随机森林作为核心分类算法.
- 用人对绩效评估进行十倍交叉验证.
主要成果:
- 在PRMxAI的预测中,单甲基氨酸的准确度为87.17%,二甲基氨酸的准确度为90.40%.
- 使用可解释的人工智能 (AI) 分析了特征的重要性.
- 随机森林分类器表现出卓越的性能.
结论:
- PRMxAI有效地预测了阿尔金因甲基化位点.
- 可解释的人工智能证实了模型的预测机制.
- PRMxAI的性能优于现有的最先进的预测工具.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Ligand Binding Sites
12.9K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
12.9K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Amino acids
89.1K
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible...
89.1K
tRNA Activation
19.3K
Aminoacyl-tRNA synthetases are present in both eukaryotes and bacteria. Though eukaryotes have 20 different aminoacyl-tRNA synthetases to couple to 20 amino acids, many bacteria do not have genes for all of these aminoacyl-tRNA synthetases. Despite this, they still use all 20 amino acids to synthesize their proteins. For instance, some bacteria do not have the gene encoding the enzyme that couples glutamine with its partner tRNA. In these organisms, one enzyme adds glutamic acid to all of the...
19.3K


