Related Experiment Video
Updated: Jan 11, 2026

Curation of Computational Chemical Libraries Demonstrated with Alpha-Amino Acids
Published on: April 13, 2022
Beyond performance: how design choices shape chemical language models
Inken Fender1,2, Jannik Adrian Gut1,2, Thomas Lemmin3
1Institute of Biochemistry and Molecular Medicine, University of Bern, Bühlstrasse 28, 3012, Bern, Switzerland.
Abstract:
Chemical language models (CLMs) have shown strong performance in molecular property prediction and generation tasks. However, the impact of design choices, such as molecular representation format, tokenization strategy, and model architecture, on both performance and chemical interpretability remains underexplored. In this study, we systematically evaluate how these factors influence CLM performance and chemical understanding. We evaluated models through fine-tuning on downstream tasks and probing the structure of their latent spaces using probing predictors, vector operations, and dimensionality reduction techniques. Although downstream task performance was similar across model configurations, substantial differences were observed in the structure and interpretability of internal representations, highlighting that design choices meaningfully shape how chemical information is encoded. In practice, atomwise tokenization generally improved interpretability, and a RoBERTa-based model with SMILES input remains a reliable starting point for standard prediction tasks, as no alternative consistently outperformed it. These results provide guidance for the development of more chemically grounded and interpretable CLMs.
Related Concept Videos
Molecular Models
Language and Cognition
Polymer Classification: Stereospecificity
Molecular Shapes
Two regions of electron density in a diatomic...
Chemical Symbols
Some symbols are derived from the common name of the element; others are abbreviations of the name in another language. Most symbols have one or two letters, but three-letter symbols have been used...
Components of Language

