iDLDDG:使用集成的深度学习功能预测DNA结合蛋白中的误解突变导致的蛋白质稳定性变化
Xuan Yu1, Fang Ge2,3, Dong-Jun Yu3
1Department of Computer Science, City University of Hong Kong, 83 Tat Chee Ave, Kowloon Tong, Hong Kong SAR(HKG), 999077, China.
Briefings in bioinformatics
|February 13, 2026
概括
预测DNA结合蛋白中的误解突变对于理解疾病至关重要. 我们新的深度学习框架,iDLDDG,准确地区分对双链和单链DNA结合蛋白的影响,改善突变预测.
科学领域:
- 基因组学和生物信息学
- 分子生物学分子生物学
- 计算生物学 计算生物学
背景情况:
- 准确预测误解突变对蛋白质-DNA结合亲和力的影响,对于疾病机制研究和治疗开发至关重要.
- 现有的模型往往无法解释双链DNA结合蛋白 (DSB) 和单链DNA结合蛋白 (SSB) 中突变的独特特征.
研究的目的:
- 开发一个计算框架,准确预测误解突变对DSB和SSB蛋白质-DNA结合亲和力的影响.
- 建立一种方法,严格区分DSB和SSB之间的突变机制.
主要方法:
- 从各种来源构建了一个全面的数据集.
- 开发了iDLDDG,这是一个深度学习框架,将基于序列的嵌入 (ESM2,ProtTrans,ESM1v) 与多个尺度的结构和进化信息集成在一起.
- 采用基于的算法来识别181个最佳残留物来建模生物物理约束,提高预测准确性和效率.
主要成果:
- iDLDDG实现了最先进的性能,在MPD276数据集上获得了0.755的10倍交叉验证皮尔森相关系数 (PCC).
- 在覆盖DSB和SSB的独立测试集上实现了0.632的PCC,显著优于现有方法.
- 证明了框架在DSB和SSB之间区分突变机制的能力.
结论:
- iDLDDG为DNA结合蛋白中病理突变的高精度预测提供了基础.
- 这项工作建立了第一个能够严格区分DSB和SSB突变机制的计算框架.
- 增强的预测精度和计算效率支持大规模评估DNA结合蛋白中的突变效应.
相关概念视频
From DNA to Protein
22.6K
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
22.6K
Single-Strand DNA Binding Proteins
16.8K
For successful DNA replication, the unwinding of double-stranded DNA must be accompanied by stabilization and protection of the separated single strands of the DNA. This crucial task is performed by single-strand DNA-binding (SSB) proteins. They bind to the DNA in a sequence-independent manner, which means that the nitrogenous bases of the DNA need not be present in a specific order for binding of SSB proteins to it. The binding of SSB proteins straightens single-stranded DNA (ssDNA) and makes...
16.8K
Mutations
94.7K
Overview
94.7K
Protein Families
17.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
17.2K
Factors Affecting Protein-Drug Binding: Protein-Related Factors
583
Drug binding to proteins is a key aspect of pharmacokinetics and can influence a drug's distribution, absorption, and elimination in the body. Several factors, including the drug's physiochemical properties, protein concentration, disease states, and the number of binding sites on the protein, influence this process.
The physicochemical properties of a drug play a significant role in its ability to bind to proteins. Lipophilic drugs, which dissolve in fats, oils, and lipids, can be...
The physicochemical properties of a drug play a significant role in its ability to bind to proteins. Lipophilic drugs, which dissolve in fats, oils, and lipids, can be...
583
Covalently Linked Protein Regulators
9.7K
Proteins can undergo many types of post-translational modifications, often in response to changes in their environment. These modifications play an important role in the function and stability of these proteins. Covalently linked molecules include functional groups, such as methyl, acetyl, and phosphate groups, and also small proteins, such as ubiquitin. There are around 200 different types of covalent regulators that have been identified.
These groups modify specific amino acids in a protein....
These groups modify specific amino acids in a protein....
9.7K


