Related Experiment Video
Updated: Jun 22, 2026

14:18
A Strategy for Sensitive, Large Scale Quantitative Metabolomics
Published on: May 27, 2014
Identification of small molecule aggregators from large compound libraries by support vector machines
Hanbing Rao1, Zerong Li, Xiangyuan Li
1College of Chemistry, Sichuan University, Chengdu 610064, People's Republic of China.
Journal of Computational Chemistry
|July 2, 2009
Summary
This study developed a support vector machine (SVM) model to identify small molecule aggregators, which often cause false positives in drug screening. The SVM model effectively distinguishes aggregators from non-aggregators in large compound libraries.
Area of Science:
- Medicinal Chemistry
- Computational Chemistry
- Drug Discovery
Background:
- Small molecule aggregators non-specifically inhibit proteins, hindering therapeutic development.
- Aggregators frequently cause false hits in high-throughput screening (HTS), necessitating their removal.
- Existing computational methods for aggregator identification lack validation on large compound libraries.
Purpose of the Study:
- To develop and validate a computational model for identifying small molecule aggregators.
- To assess the model's performance in screening large compound libraries and minimizing false positives.
Main Methods:
- Developed a support vector machine (SVM) model using 1319 known aggregators and 128,325 non-aggregators.
- Validated the SVM model using five-fold cross-validation, independent testing, and retrospective screening of large databases (13M PUBCHEM, 168K MDDR).
- Compared SVM performance against other machine learning methods.
Main Results:
- The SVM model achieved comparable aggregator and improved non-aggregator identification rates in cross-validation.
- Successfully identified 71% of independently discovered aggregators.
- Retrospective screening predicted >97% of compounds in large databases as non-aggregators, with only 1.14% of similar compounds predicted as aggregators.
- SVM demonstrated superior performance compared to other machine learning methods.
Conclusions:
- The developed SVM model is capable of accurately identifying small molecule aggregators from large compound libraries.
- The model shows substantial potential for reducing false hits in high-throughput screening campaigns.
- The findings support the use of computational methods for efficient and accurate aggregator identification in drug discovery.

