Related Experiment Video
Updated: Jan 7, 2026

Single-stage Dynamic Reanimation of the Smile in Irreversible Facial Paralysis by Free Functional Muscle Transfer
Published on: March 1, 2015
Optimizing SMILES token sequences via trie-based refinement and transition graph filtering
Sridhar Radhakrishnan1, Krish Mody2, Arvind Venkatesh3
1School of Computer Science, University of Oklahoma, Norman, OK, 73019, USA. sridhar@ou.edu.
None:
Tokenization plays a critical role in preparing SMILES strings for molecular foundation models. Poor token units can fragment chemically meaningful substructures, inflate sequence length, and hinder model learning and interpretability. Existing approaches such as SMILES Pair Encoding (SPE) and Atom Pair Encoding (APE) compress token sequences but often ignore domain-specific chemistry or fail to generalize to larger or more diverse molecules. We propose a domain-aware method for SMILES compression that combines frequency-guided substring mining using a prefix trie with an optional entropy-based refinement step using a token transition graph (TTG). On a corpus of 100,000 PubChem molecules, the Trie+TTG method reduces token sequences by more than 50% compared to APE while preserving chemically coherent substructures. The method generalizes effectively to large, out-of-distribution molecules, achieving compression rates of up to 90% with minimal sensitivity to molecule size. To assess downstream utility, we evaluate latent-space structure using unsupervised clustering and perform QSAR regression on ESOL. Trie+TTG produces more separable molecular representations and stronger predictive performance than Trie-only and APE. In addition, on peptide corpora, our method substantially outperforms SPE and the PeptideCLM tokenizer in compression and entropy metrics. These results show that combining trie-based mining with TTG refinement yields compact, stable, and chemically meaningful tokenizations suitable for modern molecular representation learning.Scientific contributions: We present a trie-based framework that compresses SMILES sequences into shorter, chemically coherent units while guaranteeing lossless reconstruction. By incorporating a token transition graph for entropy-guided refinement, our method selects contextually stable merges that improve both compression efficiency and generalization. Unlike prior approaches such as APE and SPE, our tokenizer combines frequency and context awareness, yielding more compact, interpretable, and transferable molecular representations.
Related Concept Videos
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Improving Translational Accuracy
Improving Translational Accuracy
Sequence Networks of Rotating Machines
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Heuristics
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
Rate-Determining Steps
In a multistep reaction mechanism, one of the elementary steps progresses significantly slower than the others. This slowest step is called the rate-limiting step (or rate-determining step). A reaction cannot proceed faster than its slowest step, and hence, the rate-determining step limits the overall reaction rate.
The concept of rate-determining step can be understood from the analogy of a 4-lane freeway with a short-stretch of traffic-bottleneck caused due to...
