Related Experiment Video
Updated: Jan 10, 2026

Pharmacophore Modeling for Targets with Extensive Ligand Libraries: A Case Study on SARS-CoV-2 Mpro
Published on: September 26, 2025
Ligand-based prediction of anti-bacterial compounds: Overcoming class imbalance in molecular data
Yiheng Du1, Khandaker Asif Ahmed2, Himadri Shekhar Mondal3
1College of Science and Medicine, The Australian National University, Canberra, Australia.
Abstract:
In the emergence of pan-drug-resistant bacteria, there is an urgent demand for the discovery of structurally novel antibiotics. While traditional drug development and screening procedures are cumbersome, machine learning models (ML) have shown potential to predict anti-bacterial properties of chemical compounds and shortening the screening process immensely. However, the performance of existing ML models often gets compromised, due to poorly imbalance data of vast inactive compounds, compared to handful active ones. Current study aims to leverage this class imbalance issue by introducing selective similarity-based methods for classical and neural network models, and train and test their performances across seven different models. Utilizing a combined dataset of 14,393 chemically diverse compounds, including 534 antibiotics, we utilized K-means clustering and graph-similarity measures for classical ML models (Random Forest, Logistic regression, Support vector machine, Decision tree, XGBoost and KNN) and Graph Convolutional Network (GCN) model. Further, we trained, tested and cross-validated these models and compared their performances by six different evaluation matrics. Overall, models developed utilizing our under-sampling datasets showed a significant performance increase, compared to models developed on Raw, random-sampling and randomly selected subset data. The GCN model performed best, achieving 0.97 ROC-AUC and 0.98 PRC-AUC scores. Across classical ML model, Random Forrest and Decision tree demonstrated best (ROC-AUC of 0.88 and PRC-AUC of 0.80) and worst performances (ROC-AUC of 0.73 and PRC-AUC of 0.67). We have compared our results with fourteen other similar studies, which indicate both of our results outperformed others by 10%-18%. We have discussed our findings in future drug development and screening. The code is publicly available at https://github.com/DeweyYihengDu/AbxSimU-AI.
Related Concept Videos
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Modern Molecular Taxonomy
Antibiotic Selection
The Equilibrium Binding Constant and Binding Strength
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Applications of Molecular Taxonomy

