Related Experiment Video
Updated: Feb 6, 2026

Facilitating Drug Discovery: An Automated High-content Inflammation Assay in Zebrafish
Published on: July 16, 2012
Comparing massively-multitask regression algorithms for drug discovery
Eric J Martin1, Xiang-Wei Zhu2, Patrick Riley3
1Novartis Biomedical Research, Emeryville, CA, 94608, USA. eric.martin@novartis.com.
Abstract:
Massively-multitask regression models (MMRMs) have revolutionized activity prediction for drug discovery. MMRMs trained on millions of compounds and many thousands of assays can predict bioactivity with accuracy comparable to 4-concentration IC50 experiments. This report compares six MMRMs: pQSAR, Alchemite, MT-DNN, MetaNN, Macau and IMC. Models were trained by experts in each method, on identical sets of 159 kinase and 4276 diverse ChEMBL assays, employing realistically novel training/test set splits. Results were compared both qualitatively and with statistical rigor. Our use-case is imputing full bioactivity profiles for the very sparse compound collections on which the models were trained. MMRMs performed much better than the single-task random forest regression (ST-RFR) model. Five MMRMs train all models simultaneously, so must leave out test-set measurements from all assays to avoid leakage (here 25% of data), whereas one method trains models one-at-a-time, so only holds out test data for that assay (< 1% of data). Thus, all algorithms were compared both using 75/25 splits, and when possible, 99 + / < 1 splits. Many MMRM evaluations achieved similar accuracy when tested on the same split. However, when evaluated on 75/25 splits, all MMRMs performed much worse than when evaluated on 99 + / < 1% splits. Thus, while many MMRMs produce comparable final production models (trained on all the data), models that require 75/25 splits greatly underestimate the accuracy of the final models. While outstanding for imputations, MMRMs proved little better than ST-RFR for compounds very unlike the training collection. Thus, MMRMs are best for hit-finding, off-target, promiscuity, MoA, polypharmacology or drug-repurposing within the training collection. Since accuracy is not a deciding factor, other pros and cons of each method are also described.
More Related Videos
08:49Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
05:58Using Rapid Serial Visual Presentation to Measure Set-Specific Capture, a Consequence of Distraction While Multitasking
Published on: August 29, 2018
Related Concept Videos
Regression Toward the Mean
Drug Discovery: Overview
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Correlation and Regression
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Microsoft Excel: Regression Analysis
To perform regression...