Related Experiment Video
Updated: Aug 8, 2025

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Comparative Studies on Resampling Techniques in Machine Learning and Deep Learning Models for Drug-Target Interaction
Azwaar Khan Azlim Khan1, Nurul Hashimah Ahamed Hassain Malim1
1School of Computer Sciences, Universiti Sains Malaysia, Pulau Pinang 11800, Malaysia.
Resampling techniques like SVM-SMOTE improve drug-target interaction prediction for imbalanced datasets. Deep learning models, such as Multilayer Perceptron, show high performance even without resampling.
Area of Science:
- Computational chemistry and cheminformatics
- Bioinformatics and computational biology
- Machine learning in drug discovery
Background:
- Predicting drug-target interactions (DTIs) is crucial for efficient drug discovery.
- Machine learning (ML) and deep learning (DL) methods are increasingly vital for DTI prediction.
- High-dimensional and imbalanced datasets pose significant challenges for ML/DL model training.
Purpose of the Study:
- To compare various data resampling techniques for addressing class imbalance in DTI prediction.
- To evaluate the effectiveness of deep learning methods in overcoming class imbalance for DTI prediction.
- To analyze DTI prediction performance across ten cancer-related activity classes from the BindingDB dataset.
Main Methods:
- Comparative analysis of data resampling techniques, including Random Undersampling (RUS) and SVM-SMOTE.
- Application of machine learning classifiers (Random Forest, Gaussian Naïve Bayes) and a deep learning model (Multilayer Perceptron).
- Binary classification evaluation of model performance on imbalanced DTI datasets.
Main Results:
- Random Undersampling (RUS) significantly degrades model performance on imbalanced DTI datasets, proving unreliable.
- SVM-SMOTE, when combined with Random Forest and Gaussian Naïve Bayes, achieves high F1 scores for imbalanced activity classes.
- The Multilayer Perceptron deep learning model demonstrates high F1 scores across all activity classes, irrespective of resampling.
Conclusions:
- SVM-SMOTE is a recommended resampling method for ML-based DTI prediction on imbalanced data.
- Deep learning models, particularly Multilayer Perceptron, offer robust performance in DTI prediction without requiring data resampling.
- The findings guide the selection of appropriate methods for building accurate DTI prediction models in drug discovery.
Related Concept Videos
Analysis of Population Pharmacokinetic Data
Quantitative Aspects of Drug-Receptor Interaction
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Drug Discovery: Overview

