在DNA编码的图书馆数据中普遍存在假负:链接效应如何损害基于机器学习的领先预测
Alba L Montoya1, Adam S Hogendorf1, Steven Tingey2
1Department of Medicinal Chemistry, College of Pharmacy, University of Utah 30 S 2000 E Salt Lake City UT 84112 USA raphael.franzini@utah.edu.
Chemical science
|May 21, 2025
概括
用DNA编码的化学图书馆 (DECLs) 经常会因为假负结果而错过活性药物化合物. DNA链接器影响检测,影响机器学习模型在药物发现中的准确性.
科学领域:
- 药物发现 药物发现 药物发现
- 药用化学 医学化学
- 计算化学的计算化学
背景情况:
- DNA编码的化学库 (DECL) 对于在早期药物发现中识别生物活性分子至关重要.
- DECLs生成大型数据集用于机器学习 (ML) 模型开发.
- 在DECL选择数据中的信息内容和潜在偏差尚未完全理解.
研究的目的:
- 系统地调查DECL选择中虚假负的流行情况.
- 为了确定DNA结合链接器对活性化合物的检测的影响.
- 评估DECL数据偏差对ML模型药物发现性能的影响.
主要方法:
- 作为模型系统,使用了针对PARP1/2和TNKS1/2酶的聚焦DECL.
- 分析了DECL选择数据,以量化假阴性结果并识别未被检测到的活性化合物.
- 评估了DNA结合链接器对化合物检测的影响.
- 应用低采样和超采样技术,以使用PARP2数据评估ML模型的性能.
主要成果:
- DECL选择经常产生虚假阴性,缺少许多活性化合物.
- 鉴定出DNA结合链接器是导致活性分子不足检测的一个因素.
- 假负证损害了DECL数据的预测能力,用于命中优先级和ML模型训练.
- 链接器还可以识别有针对性的选择性蛋白质参与者.
结论:
- DECL数据含有显著的偏差,特别是假阴性,影响其在药物发现中的有用性.
- 在DECL中,DNA结合链接器既带来了挑战 (低检测),也带来了机遇 (选择性识别).
- 数据处理和ML模型开发的最佳实践对于最大限度地提高DECL数据的价值至关重要.
相关概念视频
Mismatch Repair
38.1K
Overview
38.1K
The Central Dogma
116.7K
Overview
116.7K
Nonsense-mediated mRNA Decay
9.4K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
9.4K
Genome Copying Errors
4.3K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.3K
The Central Dogma
21.2K
The central dogma explains the flow of genetic information from DNA nucleotides to the amino acid sequence of proteins.
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
RNA is the Missing Link Between DNA and Proteins
In the early 1900s, scientists discovered that DNA stores all the information needed for cellular functions and that proteins perform most of these functions. However, the mechanisms of converting genetic information into functional proteins remained unknown for many years. Initially, it was believed that a single gene is...
21.2K
Mismatch Repair
5.4K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
5.4K


