Related Experiment Video
Updated: Jun 12, 2026

DNA Nanotubes as a Versatile Tool to Study Semiflexible Polymers
Published on: October 25, 2017
A Canonical Text Representation for Polymers via BigSMILES and Tree Automata
Bruno S Leão1, Nathan J Rebello1, Bradley D Olsen1
1Department of Chemical Engineering, Massachusetts Institute of Technology, 77 Massachusetts Avenue, Cambridge, Massachusetts 02139, United States.
None:
A BigSMILES string encodes the structural connectivity of any polymer chemistry and topology as a linear string. However, multiple BigSMILES strings can encode the same ensemble, making string-based searches for polymers in digital databases challenging. This work presents a canonicalization algorithm that breaks the degeneracy of the BigSMILES language for both linear and branched polymers and can reverse-translate canonicalized structures back into BigSMILES. The algorithm was validated on broadly representative polymer chemistries and topologies from the literature. First, the BigSMILES string is mapped onto a tree automaton, a type of state machine that accommodates branch points and recognizes the same ensemble of molecules that BigSMILES encodes. The automaton can then be minimized into a unique graph with the fewest states through existing algorithms. Finally, a human-readable canonicalized BigSMILES is obtained upon translation of the state machine transition rules back into a string. This robust canonicalization algorithm allows polymers to be searched rapidly in large database systems, making data findable, accessible, interoperable, and reusable (FAIR) and enabling the development of novel data-driven approaches with BigSMILES.
Related Concept Videos
Polymers
Polymers
Polymers
Characteristics and Nomenclature of Copolymers
Ziegler–Natta Chain-Growth Polymerization: Overview
Characteristics and Nomenclature of Homopolymers

