Related Experiment Video
Updated: Sep 30, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Chemical language models for early-stage drug discovery: applications, pitfalls, and future directions
1College of Pharmacy, Al Ain University, P.O. Box 112612, Abu Dhabi, United Arab Emirates. Abdallah.abouhajal@gmail.com.
Abstract:
Attrition rates in clinical drug development remain above 90%, with most failures rooted in inadequate compound quality at the early discovery stage. Chemical language models (CLMs), neural sequence models pretrained on hundreds of millions of SMILES or SELFIES strings under self-supervised language-modeling objectives, have emerged as a versatile tool class across the early-stage pipeline. This review covers their foundations, including molecular representation, tokenization, and recurrent, encoder, decoder, and encoder-decoder families, before examining three application domains: virtual screening, property prediction, and de novo generation. In virtual screening, CLM embeddings retrieved by approximate nearest-neighbor search offer sub-second triage of billion-scale compound libraries, although reported evaluations remain largely retrospective, and pairing CLMs with protein language models extends this to target-aware ranking at comparable scale. In property prediction, sequence-only transformers including ChemBERTa and MolBERT are competitive on ADMET endpoints without consistently surpassing well-tuned descriptor-based and graph-based baselines, and hybrid architectures fusing sequence embeddings with structural priors represent the most promising direction. In de novo generation, reinforcement learning and scaffold-hopping strategies have advanced CLMs into prospective campaigns, with designs confirmed active in cell-based and biochemical assays across nuclear-receptor and lipid-kinase targets. Persistent limitations include poor interpretability and attendant regulatory uncertainty, performance brittleness under distribution shift without calibrated uncertainty estimates, and computational demands concentrated in well-resourced laboratories. Addressing these constraints may depend less on further architectural refinement than on multimodal foundation models, calibrated uncertainty quantification reported as first-class inference outputs, mandatory scaffold-split evaluation protocols, and integration into agentic closed-loop workflows coupling prediction with physics-based simulation and laboratory automation.
Related Concept Videos
Drug Discovery: Overview
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence its...
Pharmacodynamic Models: Overview
Molecular Models
Pharmacogenomics: Identification of New Drug Targets
Principles of Drug Action
Drugs can be agonists or antagonists. Like the endogenous ligands, agonists always bind and activate the target to produce a cellular response. Agonist binding induces a conformational change which in turn...
