预测DRBP-MLP:通过多层感知子预测DNA结合蛋白和RNA结合蛋白
Ozgur Can Arican1, Ozgur Gumus2
1Department of Health Bioinformatics, Ege University, 35100, Izmir, Turkey.
Computers in biology and medicine
|August 10, 2023
概括
一个新的多层感知子 (MLP) 模型,PredDRBP-MLP,有效地分类DNA结合蛋白 (DBP),RNA结合蛋白 (RBP) 和非核酸结合蛋白 (NNABP). 这种模型需要更少的计算能力,并且比CNN-BiLSTM方法更快.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习在生物学中的应用
背景情况:
- 与DNA (DNA结合蛋白,DBP) 和RNA (RNA结合蛋白,RBP) 相互作用的蛋白质对于细胞功能至关重要.
- 准确的DBPs,RBPs和非核酸结合蛋白 (NNABP) 的分类对于理解生物过程至关重要.
- 现有的CNN-BiLSTM模型用于DBP/RBP分类是计算密集且耗时的.
研究的目的:
- 开发一个高效和准确的人工学习模型来分类DBPs,RBPs和NNABPs.
- 引入基于多层感知子 (MLP) 的预测器PredDRBP-MLP,作为现有方法的替代方案.
- 为蛋白质分类提供更快,更少资源密集的解决方案.
主要方法:
- 开发PredDRBP-MLP,这是一个使用多层感知器 (MLP) 架构的人工学习模型.
- 实施DBP,RBP和NNABP的多类分类方法.
- 在独立数据集上评估PredDRBP-MLP性能,并将其与现有预测器进行比较.
主要成果:
- 预测DRBP-MLP表现出成功的分类,与其他预测因素相比,在NNABP类中表现出色.
- 在NNABP类中,PredDRBP-MLP实现了0.578的精度,0.522的回忆和0.549的F1得分.
- 基于MLP的模型需要较低的处理能力,并且比CNN-BiLSTM模型快得多.
- 为PredDRBP-MLP开发了一个桌面应用程序,可以自由访问.
结论:
- PredDRBP-MLP提供了一种计算高效和有效的方法来分类DNA结合蛋白,RNA结合蛋白和非核酸结合蛋白.
- 开发的MLP模型为生物信息学研究提供了有价值的工具,提供了更好的性能和速度.
- 由于PredDRBP-MLP桌面应用程序的可访问性,使其在生物研究中更容易使用.
相关概念视频
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Cooperative Binding of Transcription Regulators
6.5K
Transcriptional regulators bind to specific cis-regulatory sequences in the DNA to regulate gene transcription. These cis-regulatory sequences are very short, usually less than ten nucleotide pairs in length. The short length means that there is a high probability of the exact same sequence randomly occurring throughout the genome. Since regulators can also bind to groups of similar sequences, this further increases the chances of random binding. Transcriptional regulators form...
6.5K
RNA Polymerase II Accessory Proteins
9.2K
Proteins that regulate transcription can do so either via direct contact with RNA Polymerase or through indirect interactions facilitated by adaptors, mediators, histone-modifying proteins, and nucleosome remodelers. Direct interactions to activate transcription is seen in bacteria as well as in some eukaryotic genes. In these cases, upstream activation sequences are adjacent to the promoters, and the activator proteins interact directly with the transcriptional machinery. For example, in...
9.2K
lncRNA - Long Non-coding RNAs
8.6K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.6K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K


