Related Experiment Video
Updated: Mar 28, 2026

Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
Multitask Pretraining Framework for Improving Predictivity of Machine Learning Chemical Bioactivity Models for
Noah J Wichrowski1, Mary Versa Clemens-Sewall1, Karun K Rao1
1The Johns Hopkins University Applied Physics Laboratory, Laurel, Maryland 20723, United States.
None:
Computational models are crucial for rapid hazard screening of novel chemicals when time and resources are not available for laboratory assessment. The rise of machine learning (ML) methods powering quantitative structure-activity relationship (QSAR) models has enabled data-driven development of predictive models for health effects screening. However, these models are typically single-task, meaning that they are trained on a single toxicological endpoint and lack transferability to similar tasks, i.e., the ability to predict chemicals' effects on related endpoints. Thus, when predictions are needed for another endpoint, a new model must be trained from scratch. Further, single-task ML models are typically trained on very large, homogeneous data sets, which are not available for most adverse outcome endpoints. Effective hazard screening would benefit from approaches that can handle multiple small, noisy data sets recording complex chemical and biological mechanisms. To that end, we trained an ML model simultaneously on multiple tasks curated from moderate-sized (∼1000 observations) ToxCast data sets. To predict novel tasks from small (∼100 observations) ToxCast data sets, we combined our pretrained multitask model with a task-specific predictor, either a random forest or a neural network. These two components comprise a novel ML pipeline that generates and uses molecular representations from our multitask model. Compared to a common ML approach using standard chemical representations, our pipeline performed statistically better on a majority of tasks, regardless of the choice of downstream predictor. The advantage of the molecular representations from our multitask model, over those from a single-task model, is that they combine information on multiple effects to provide a model of chemical space that captures generalizable information. This work contributes to efforts to improve the utility of ML QSAR methods for predicting chemicals' bioactivity on low-data toxicological endpoints.
Related Concept Videos
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Drug Discovery: Overview
Model Approaches for Pharmacokinetic Data: Physiological Models
Preclinical Development: Overview
Predicting Reaction Outcomes

