Data mining a small molecule drug screening representative subset from NIH PubChem

Xiang-Qun Xie1, Jian-Zhong Chen

  • 1Department of Pharmaceutical Sciences, School of Pharmacy, Pittsburgh Molecular Library Screening Center, Drug Discovery Institute, Pittsburgh, Pennsylvania 15260, USA. xix15@pitt.edu

Summary

Researchers developed a data mining approach to create rePubChem, a diverse 540K compound subset from PubChem, ideal for efficient drug screening. This method ensures structural diversity and basic properties are maintained for virtual and high-throughput screening.