Related Experiment Videos
Database clustering with a combination of fingerprint and maximum common substructure methods
1F. Hoffmann-La Roche AG, Pharmaceutical Research, CH-4070 Basel, Switzerland.
Journal of Chemical Information and Modeling
|June 1, 2005
Summary
This study introduces an efficient stepwise method for clustering large chemical databases using exclusion sphere and maximum common substructure algorithms. It effectively identifies conserved scaffolds and prioritizes chemical libraries.
Area of Science:
- Cheminformatics
- Computational Chemistry
- Drug Discovery
Background:
- Clustering large chemical databases is crucial for scaffold identification and library prioritization.
- Existing methods may lack efficiency or precision in handling vast datasets.
Purpose of the Study:
- To develop an efficient, stepwise method for clustering large chemical databases.
- To identify conserved substructures and distinct chemical entities within databases.
Main Methods:
- Stepwise clustering using an extended exclusion sphere algorithm with Tanimoto coefficients and Daylight fingerprints.
- Iterative extraction of maximum common substructures from clusters.
- Merging of clusters based on common substructures and similarity to singletons.
Main Results:
- The method successfully identifies tight clusters with conserved substructures.
- It generates singletons only for truly distinct chemical structures.
- Demonstrated utility in identifying frequent scaffolds, selecting analogues, and prioritizing commercial libraries.
Conclusions:
- The developed method provides an efficient and accurate approach for chemical database clustering.
- It aids in understanding chemical space, facilitating drug discovery and library management.