Related Experiment Video
Updated: Aug 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Diversity Beats Size Scaling for Chemical Language Models
Borja Medina1, Alessandro Tibo1, Jiazhen He1
1Molecular AI, Discovery Sciences, R&D, AstraZeneca, Gothenburg, Sweden.
Abstract:
Chemical language models, such as transformers trained on SMILES strings, are increasingly used in drug design and have seen rapid growth in both model capacity and training dataset size. The impact of this scaling on practical downstream performance remains unclear, however. We systematically evaluate how model size and dataset size affect encoder-decoder transformers trained on paired textual molecular representations. We find that, beyond a minimal threshold, further model scaling yields no gain in hit generation rate, while dataset scaling gives diminishing returns. We further introduce a dataset diversification strategy that substantially increases hit diversity. These results suggest that, for molecular hit discovery, data curation and diversity may be more impactful than continued scaling of model size or dataset volume, and they motivate a shift from scale-first to diversity-first training paradigms.
Related Concept Videos
Scaling
Introduction to Scalers
Scalar...
One-Way ANOVA: Unequal Sample Sizes
Scale-Up Processes
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...