对比大规模多任务回归算法用于药物发现
Eric J Martin1, Xiang-Wei Zhu2, Patrick Riley3
1Novartis Biomedical Research, Emeryville, CA, 94608, USA. eric.martin@novartis.com.
Journal of computer-aided molecular design
|February 4, 2026
概括
与单任务模型相比,大规模多任务回归模型 (MMRMs) 显著改善了药物发现活动预测. 然而,当使用大型测试集时,它们的性能被高估了,这凸显了现实数据分割对于准确评估的重要性.
科学领域:
- 计算化学计算化学
- 化学信息学 化学信息学
- 机器学习在药物发现中的作用
背景情况:
- 大规模多任务回归模型 (MMRMs) 已经成为预测药物发现中的化合物生物活性的强大工具.
- 这些模型在广泛的数据集上进行训练,提供与实验测量可比的准确性.
研究的目的:
- 为了比较六个主要的MMRM (pQSAR,Alchemite,MT-DNN,MetaNN,澳门,IMC) 在生物活性概况归算方面的表现.
- 评估不同培训/测试组分离对MMRM性能和准确性估计的影响.
主要方法:
- 六名MMRM受过专家对相同的数据集进行培训,包括159个酶和4276个ChEMBL测定.
- 模型使用75/25和99+/<1%的训练/测试集分割进行评估,以评估在不同数据可用性场景下的性能.
- 对比分析包括定性评估和统计严谨性,与单任务随机森林回归 (ST-RFR) 进行基准测试.
主要成果:
- 在生物活性概况归算方面,MMRM显著优于ST-RFR模型.
- 根据训练/测试分割,性能差异很大;与99+/<1%分割相比,75/25分割导致模型准确性的大幅低估.
- 虽然MMRMs擅长在训练数据分布中归因配置,但对于与训练集不同的化合物,它们的优势会减少.
结论:
- 在训练数据的化学空间内,MMRM对于诸如命中发现,目标外预测和药物再利用等任务非常有效.
- 数据分割策略的选择对MMRM的感知准确性产生了重大影响,需要使用现实的,较小的测试集来进行可靠的评估.
- 对于探索已知的化学空间而言,MMRM的实用性最大,而它们对新型化学实体的性能则需要进一步调查.
更多相关视频
08:49Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
1.3K
05:58Using Rapid Serial Visual Presentation to Measure Set-Specific Capture, a Consequence of Distraction While Multitasking
Published on: August 29, 2018
9.3K
相关概念视频
Regression Toward the Mean
7.1K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.1K
Drug Discovery: Overview
11.7K
Drug discovery is a multifaceted process involving extensive screening, testing, and optimization of lead compounds to identify potential new drugs for therapeutic use. It combines several approaches, including screening large numbers of natural products, chemical modification of known active molecules, identification of new drug targets, and rational design based on biological mechanisms and drug-receptor structure. These approaches are carried out in both academic research laboratories and...
11.7K
Multiple Regression
4.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.0K
Correlation and Regression
3.5K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
3.5K
Regression Analysis
8.4K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.4K
Microsoft Excel: Regression Analysis
1.6K
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
1.6K
