Related Experiment Video
Updated: Jun 25, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Balancing Data Quantity and Quality: Evaluating Curation Strategies for Bioactivity Prediction in Lead Optimization.
Carl C G Schiebroek1, Gregory A Landrum1, Sereina Riniker1
1Department of Chemistry and Applied Biosciences, ETH Zurich, Vladimir-Prelog-Weg 2, 8093 Zurich, Switzerland.
Developing accurate machine-learning (ML) models for predicting chemical bioactivity is difficult. Our study found that increasing data quantity, even with noise, did not improve ML model generalization for lead optimization.
Area of Science:
- Computational chemistry
- Cheminformatics
- Machine learning in drug discovery
Background:
- Accurate machine-learning (ML) models for predicting chemical bioactivity require large, diverse, and low-noise training datasets.
- Public databases like ChEMBL may contain data with varying curation rigor, impacting dataset size, diversity, and noise levels.
- The trade-off between increasing dataset size and introducing noise is not well understood for model generalization.
Purpose of the Study:
- To compare three data curation and modeling strategies for predicting chemical bioactivity.
- To assess the impact of data quantity versus label consistency on model generalization.
- To evaluate the effectiveness of multitask learning (MTL) and graph neural networks (GNNs) in bioactivity prediction.
Main Methods:
- Compared models trained on single-target data, single-assay condition data, and multitask learning (MTL) models.
- Utilized graph neural networks (GNNs) and random forests (RF) regressors.
- Employed a leave-assay-out cross-validation strategy to minimize noise in test sets.
Main Results:
- No significant performance differences were observed between the three data curation strategies.
- Increasing data quantity at the expense of label consistency did not improve model generalization for lead optimization tasks.
- The MTL approach did not offer a performance advantage over simpler methods.
- Graph neural networks (GNNs) showed high seed-dependent variability, necessitating multi-seed evaluation for reliable assessment.
Conclusions:
- For lead optimization, enhancing data quantity without ensuring label consistency does not necessarily improve machine-learning model generalization.
- Multitask learning (MTL) did not outperform other strategies in this context.
- Robust model assessment for bioactivity prediction requires careful consideration of seed variability, especially with limited training data.
Related Concept Videos
Drug Discovery: Overview
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence its...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
The Equilibrium Binding Constant and Binding Strength
Bioequivalence Data: Statistical Interpretation
Bioavailability Study Design: Healthy Subjects Versus Patients