DBPMod:一种监督学习模型,用于模型生物中的DNA结合蛋白的计算识别
Upendra K Pradhan1, Prabina K Meher1, Sanchita Naha2
1Division of Statistical Genetics, ICAR-Indian Agricultural Statistics Research Institute, PUSA, New Delhi 110012, India.
Briefings in functional genomics
|August 31, 2023
概括
一种新的计算方法,DBPMod,使用机器学习和进化特征准确识别特定物种的DNA结合蛋白 (DBPs). 这种工具超越了现有的方法,有助于理解关键的生物过程.
科学领域:
- 分子生物学分子生物学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 基因结合蛋白 (DBP) 对于基因调节和DNA修复等基本生物过程至关重要.
- 准确识别DBP对于理解分子机制至关重要.
- 现有的计算方法往往缺乏识别特定物种DBP的特异性.
研究的目的:
- 开发一种新的计算方法,DBPMod,用于准确识别特定物种的DNA结合蛋白.
- 通过利用特定物种的特征和机器学习模型来提高DBP预测的准确性.
主要方法:
- 开发了基于机器学习的计算方法DBPMod.
- 采用浅层和深度学习算法进行预测.
- 利用进化特征和序列衍生特征用于模型训练.
- 在五种模型生物体中验证了性能: *C. elegans*, *D. melanogaster*, *E. coli*, *H. sapiens* 和 *M. musculus*.
- 使用五倍交叉验证和独立测试集评估准确性,在接收器操作特征曲线 (auROC) 下的测量面积和在精度回忆曲线 (auPRC) 下的测量面积.
主要成果:
- 与深度学习模型相比,浅层学习模型的准确性更高.
- 进化特征在预测准确性方面被证明比序列衍生特征更有效.
- DBPMod实现了高的预测准确度,auROC在模型生物体中大约在89-92%之间,auPRC在89-95%之间.
- 对于所有测试的物种,DBPMod在DBP识别方面表现优于现有的12种最先进的计算方法.
- 为DBPMod开发了一个公开可访问的Web服务器.
结论:
- DBPMod提供了一个高度准确和特定物种的方法来识别DNA结合蛋白.
- 该方法依赖于进化特征,增强了其预测能力.
- 对于研究人员来说,DBPMod是一个有价值的工具,它补充了实验和计算DBP发现工作.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Single-Strand DNA Binding Proteins
14.2K
For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
14.2K
Labeling DNA Probes
8.2K
DNA probes are fragments of DNA labeled with a reporter tag to enable their detection or purification. The resulting labeled DNA probes can then hybridize to target nucleic acid sequences through complementary base-pairing, and may be used to recover or identify these regions.
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
8.2K
Cooperative Binding of Transcription Regulators
6.5K
Transcriptional regulators bind to specific cis-regulatory sequences in the DNA to regulate gene transcription. These cis-regulatory sequences are very short, usually less than ten nucleotide pairs in length. The short length means that there is a high probability of the exact same sequence randomly occurring throughout the genome. Since regulators can also bind to groups of similar sequences, this further increases the chances of random binding. Transcriptional regulators form...
6.5K


