Related Experiment Video
Updated: Dec 17, 2025

Author Spotlight: Generating Neuronal Phenotypic Profiles - A Protocol to Culture and Image Human Midbrain Dopaminergic Neurons
Published on: July 7, 2023
Deep Learning-Based Imbalanced Data Classification for Drug Discovery.
1Trakya University Faculty of Medicine, Department of Biostatistics and Medical Informatics, Edirne, Turkey.
Data balancing methods can improve deep neural network performance on imbalanced compound data from high-throughput screening (HTS) in drug discovery. However, severe data imbalance still negatively impacts classification accuracy.
Area of Science:
- Computational chemistry and cheminformatics
- Machine learning in drug discovery
- Bioassay data analysis
Background:
- Drug discovery is costly and time-consuming, with early stages focusing on identifying and optimizing drug-like compounds.
- High-throughput screening (HTS) is a conventional method for detecting active compounds, generating vast datasets.
- The PubChem repository offers millions of HTS bioassays, valuable for machine learning but often suffer from imbalanced data.
Purpose of the Study:
- To investigate the classification performance of deep neural networks (DNNs) on imbalanced compound datasets from HTS.
- To evaluate the effectiveness of various data balancing techniques in mitigating performance issues caused by data imbalance.
- To assess the impact of the degree of data imbalance on DNN classification performance.
Main Methods:
- Utilized five confirmatory HTS bioassays from the PubChem repository.
- Applied one undersampling and three oversampling methods for data balancing.
- Employed a fully connected, two-hidden-layer DNN for classifying active and inactive molecules.
- Evaluated performance using balanced accuracy, precision, recall, F1 score, Matthews correlation coefficient, and AUC.
Main Results:
- Data balancing methods demonstrated a degree of mitigation for the negative effects of imbalanced data on DNN performance.
- The extent of data imbalance was found to negatively correlate with the overall classification performance of the network.
- Specific balancing methods showed varying degrees of success in improving classification metrics for imbalanced datasets.
Conclusions:
- Data balancing techniques are crucial for improving machine learning model performance on imbalanced HTS bioassay data.
- While balancing methods help, significant data imbalance remains a challenge in accurately classifying active compounds.
- Further research is needed to optimize balancing strategies for complex, imbalanced datasets in drug discovery.
Related Concept Videos
Drug Discovery: Overview
Drug Classes and Categories
Therapeutic Drug Monitoring: Overview and Classification
Cardiovascular Drugs: Classification based on Therapeutic Indications
Classification of Neurotransmitters
Drug Distribution: Plasma Protein Binding
