扩大基于结构的受体-连接物结合亲和关系回归的训练数据,通过缺失标签的归算来进行这种回归
Paul G Francoeur1, David R Koes1
1Department of Computational and Systems Biology, University of Pittsburgh, Pittsburgh, Pennsylvania 15260, United States.
这项研究探讨了通过赋予绑定亲和标签来扩展有限的分子性质预测数据. 这种机器学习方法显示,尽管增加了噪音,但基于结构的模型的性能略有改善.
科学领域:
- 计算化学是一种计算化学.
- 机器学习 机器学习
- 药物发现 药物发现
背景情况:
- 机器学习模型需要大量的数据集才能进行有效的训练.
- 目前基于结构的分子性质预测模型面临有限的训练数据,特别是受体-连接体结合亲缘关系.
- 像CrossDocked2020这样的现有数据集扩展了绑定姿势分类数据,但没有绑定亲和力数据.
研究的目的:
- 调查赋予绑定亲和标签的可行性,以扩展基于结构的分子性质预测模型的训练数据.
- 评估使用假设标签对绑定亲缘关系回归模型性能的影响.
主要方法:
- 对于缺乏实验数据的分子复合体,使用了假定的结合亲和度标签.
- 训练了一个卷积神经网络,使用来自CrossDocked2020数据集的现有绑定亲和数据.
- 评估基于结构的模型的性能,这些模型使用实验数据和假定的结合亲和数据进行训练.
主要成果:
- 赋予绑定亲和标签是增强培训数据集的可行策略.
- 用假设标签训练的模型在结合亲和力回归性能上显示了微小的改善.
- 计入标签在培训数据中增加了额外的噪音,但性能仍然有所改善.
结论:
- 结合性亲和标签的推算提供了一种实用方法,用于增加基于结构的分子建模的训练数据.
- 尽管存在固有的噪音,但归算数据可以提高机器学习模型对受体-连接体结合亲和力的预测准确度.
- 这种方法通过克服数据限制,有助于推进计算药物发现.
更多相关视频
08:31Biosensor-based High Throughput Biopanning and Bioinformatics Analysis Strategy for the Global Validation of Drug-protein Interactions
Published on: December 1, 2020
10:29Quantitative Structure-Activity Relationship, Activity Prediction, and Molecular Dynamics of Non-nucleotide Reverse Transcriptase Inhibitors
Published on: May 9, 2025
相关概念视频
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
The Equilibrium Binding Constant and Binding Strength
Mechanistic Models: Compartment Models in Individual and Population Analysis
Quantitative Aspects of Drug-Receptor Interaction
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
