Related Experiment Video
Updated: Aug 6, 2026

Synthesis of Information-bearing Peptoids and their Sequence-directed Dynamic Covalent Self-assembly
Published on: February 6, 2020
HELMify: A Hybrid Rule- and LLM-Based Generator of Peptide Monomer HELM Names
Robert P Sheridan1, Rajvi Shah2, Kathryn McGarty2
1Modeling and Informatics, Merck & Co., Inc., Rahway, New Jersey07065, United States.
Abstract:
HELM is a hierarchical notation system for biopolymers that is an increasingly popular choice for representing peptides. In this system, each monomer name must uniquely identify a single monomer, and until now, chemists have named peptide monomers by hand using a limited set of rules. As the chemical space of monomers is practically unlimited, this naming process is too time-consuming to be done by chemists and is prone to inconsistency. To address this bottleneck, we developed a public benchmark data set for this task and two complementary automated naming approaches for HELM name generation: a zone-based namer, which relies on hand-curated dictionaries of substructures, and a generative AI namer, which relies on a pretrained Large Language Model (LLM) with chemistry world knowledge. We evaluated both approaches on a set of publicly disclosed alpha amino acids. For a time-split test set, the zone-based namer produced usable names for 84% of cases: 54% matched the human-generated name exactly, and an additional 30% did not match but were interpretable as the correct structure. The LLM-based namer produced usable names for 42% of cases: 33% of the names matched the human-generated name exactly, and an additional 9% were nonidentical but still interpretable as the correct structure. In another 49% of LLM-generated cases, the outputs were considered close matches, and while not directly usable, they can still inform the generation of correct names. The zone-based namer is deterministic, fast, and independent of any previous names, explicitly handling zone order and chirality. The LLM-based namer can generalize from human-named examples to propose names for novel chemical groups, helping fill library gaps. Together, the approaches offer complementary coverage: rules-based precision and speed paired with data-driven flexibility for novel moieties. Combining the strengths of both approaches paves the way for an automated peptide monomer naming process that can accelerate research by (i) helping the community establish standardized guidelines for the rapidly expanding noncanonical amino acid (ncAA) space, (ii) streamlining integration with synthesis planning, inventory management, and analysis pipelines, (iii) enabling more consistent annotation in ELNs, databases, and publications, and (iv) facilitating cross-lab replication and data exchange. Together, these advances can expedite method development, reduce ambiguity in collaboration, and ultimately support the discovery of novel peptides with improved drug-like properties.
More Related Videos
12:02An Efficient Method for the Synthesis of Peptoids with Mixed Lysine-type/Arginine-type Monomers and Evaluation of Their Anti-leishmanial Activity
Published on: November 2, 2016
08:55Facile Protocol for the Synthesis of Self-assembling Polyamine-based Peptide Amphiphiles (PPAs) and Related Biomaterials
Published on: June 25, 2018
Related Concept Videos
Characteristics and Nomenclature of Homopolymers
Olefin Metathesis Polymerization: Acyclic Diene Metathesis (ADMET)
Similar to cross-metathesis, ADMET also involves the formation of metallacyclobutane intermediate by [2+2] cycloaddition of one of the double bonds of a terminal diene with...
Characteristics and Nomenclature of Copolymers
Nomenclature of Carboxylic Acid Derivatives: Acid Halides, Esters, and Acid Anhydrides
The IUPAC and common names of acid halides are derived from the corresponding carboxylic acids, by changing “ic acid” to “yl halide.” For example, as shown below, the IUPAC name ethanoyl chloride is derived from ethanoic acid, and the common name, acetyl chloride, is obtained from acetic acid.
¹H NMR: Pople Notation
A proton...
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...