Related Experiment Video
Updated: Jun 28, 2026

05:12
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
Toward an improved clustering of large data sets using maximum common substructures and topological fingerprints
1Boechringer Ingelheim (Canada) Ltd. Research & Development, 2100 Cunard Street, Laval, Quebec, Canada H7S 2G5. Alexander.bocker@boehringer-ingelheim.com
Journal of Chemical Information and Modeling
|October 30, 2008
Summary
A novel clustering algorithm efficiently groups large molecular datasets by chemotype. This method enhances the analysis of high-throughput screening results and structure-activity relationships.
Area of Science:
- Computational chemistry
- Cheminformatics
- Data mining
Background:
- Analyzing large chemical datasets is crucial for drug discovery.
- Existing methods struggle with scalability and accurate chemotype identification.
- Understanding structure-activity relationships (SAR) requires robust molecular grouping.
Purpose of the Study:
- To develop a scalable algorithm for clustering large datasets based on chemotypes.
- To enable improved analysis of high-throughput screening (HTS) and virtual screening (VS) data.
- To facilitate insights into structure-activity relationships and chemotype hopping.
Main Methods:
- Hierarchical k-means algorithm applied to molecular fingerprints for preclustering.
- Maximum common substructure (MCS) approach for chemotype extraction from terminal clusters.
- Iterative fusion of similar chemotypes and singletons, followed by overlap-based grouping.
- Second round of hierarchical k-means using chemotype representatives for final hierarchical grouping.
Main Results:
- Successful grouping of large molecular datasets (>100,000 molecules) by chemotype.
- Identification of chemotypes based on shared structural features (rings, atoms, heavy atoms).
- Demonstration of utility in analyzing reverse transcriptase inhibitors and evaluating similarity searching routines.
- Interactive graphical user interface for visualizing SAR and chemotype hopping potential.
Conclusions:
- The developed algorithm provides a high-quality, scalable solution for analyzing large chemical datasets.
- It enhances the analysis of HTS and VS results, aiding in drug discovery and lead optimization.
- The method effectively identifies chemotypes and supports the exploration of chemotype hopping strategies.
Related Concept Videos
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
