Related Experiment Videos
Clustering a large number of compounds. 2. Using the Connection Machine
Summary
A new protocol allows screening of 230,000 National Cancer Institute Repository compounds. Clustering on a massively parallel computer efficiently extracts a representative sample for drug discovery.
Area of Science:
- Computational chemistry
- Cheminformatics
- Drug discovery
Background:
- The National Cancer Institute (NCI) Repository contains approximately 230,000 compounds.
- These compounds are available for screening to identify potential drug candidates.
- Efficiently sampling these large chemical libraries is crucial for drug discovery efforts.
Purpose of the Study:
- To describe a project focused on extracting a representative sample of compounds from the NCI Repository.
- To detail the application of clustering techniques for this sampling task.
- To report on the implementation of a clustering program on a massively parallel computer.
Main Methods:
- Utilized a clustering program to analyze and sample the NCI compound repository.
- Implemented the clustering algorithm on the Connection Machine, a parallel processing computer.
- Leveraged the 16,000 processing elements of the Connection Machine for computational efficiency.
Main Results:
- The clustering approach successfully facilitated the extraction of a representative compound sample.
- The implementation on the Connection Machine transformed a complex task into a manageable process.
- This method enables routine and efficient analysis of large chemical datasets.
Conclusions:
- Massively parallel computing offers a powerful solution for analyzing large chemical libraries.
- Clustering is an effective strategy for generating representative compound subsets for screening.
- This approach enhances the efficiency of drug discovery by optimizing sample selection.