Related Experiment Video
Updated: Apr 1, 2026

08:21
Curation of Computational Chemical Libraries Demonstrated with Alpha-Amino Acids
Published on: April 13, 2022
3.2K
Get Your Atoms in Order--An Open-Source Implementation of a Novel and Robust Molecular Canonicalization Algorithm
Nadine Schneider1, Roger A Sayle2, Gregory A Landrum1
1Novartis Institutes for BioMedical Research, Novartis Pharma AG , Novartis Campus, CH-4002 Basel, Switzerland.
Journal of Chemical Information and Modeling
|October 7, 2015
Summary
A new algorithm offers a robust and fast method for creating unique molecular representations by establishing a canonical atom ordering. This approach overcomes limitations of existing methods, enabling universal molecule identification.
Area of Science:
- Cheminformatics
- Computational Chemistry
- Drug Discovery
Background:
- Canonical molecular ordering is crucial for unique molecule representation.
- Existing methods like the Morgan algorithm have limitations with large molecules and inconsistencies across toolkits.
- Lack of a universal standard hinders molecule identification.
Purpose of the Study:
- To develop an alternative canonicalization approach for molecules.
- To address limitations of current graph relaxation algorithms.
- To facilitate the generation of universal unique molecule identifiers.
Main Methods:
- Utilized a standard stable-sorting algorithm instead of Morgan-like indexing.
- Developed two new invariants for handling dependent chirality and symmetrical cyclic graphs.
- Tested the algorithm on the ChEMBL 20 dataset (1.45 million compounds).
Main Results:
- The new approach demonstrated robustness and speed, even with random atom renumbering and SMILES round tripping.
- Achieved canonical ordering for protein molecules in milliseconds.
- The algorithm is implemented in the RDKit toolkit and provided with a Python reference implementation.
Conclusions:
- The novel algorithm provides a reliable and efficient method for canonical atom ordering.
- This work represents a significant step towards a common standard for universal molecule identification.
- The open-source implementation facilitates integration into various cheminformatics toolkits.
Related Concept Videos
Applications of Molecular Taxonomy
655
Molecular taxonomy has revolutionized the understanding and classification of bacteria, providing precise insights into their diversity, evolutionary relationships, and ecological roles. By utilizing molecular techniques such as DNA sequencing and fingerprinting, researchers have made significant strides in various fields related to bacterial studies.Resolving Taxonomic AmbiguitiesMolecular taxonomy has been instrumental in distinguishing closely related bacterial species initially thought to...
655
Modern Molecular Taxonomy
835
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
835
Molecular Models
45.6K
Physical models representing molecular architectures of chemical compounds play essential roles in understanding chemistry. The use of molecular models makes it easier to visualize the structures and shapes of atoms and molecules.
45.6K
Complementary DNA
32.2K
Overview
32.2K
Complementary DNA
7.3K
7.3K
Nucleic Acid Structure
10.3K
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
DNA Structure
DNA...
10.3K

