Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Non-invasive Multimodal Cardiovascular Disease Detection Method Based on Comprehensive View Analysis.

Journal of imaging informatics in medicine·2026
Same author

Plasma protein GDF15 has a good predictive potential for the kidney complications of type 2 diabetes.

Frontiers in endocrinology·2026
Same author

DAVID: a web server for functional annotation and functional enrichment analysis of gene lists (2025 update).

Nucleic acids research·2026
Same author

Interfacial anionic competition-driven electrochemical evolution in FeF<sub>3</sub> conversion electrodes.

Nature communications·2026
Same author

Body fat, eating behaviours, and well-being as predictors of negative body talk among college students: a moderation analysis by sex.

BMC public health·2026
Same author

Spatial distribution, contamination characteristics and health hazard potential of soil potentially toxic elements under different reclamation modes in coal mining subsidence areas.

Environmental monitoring and assessment·2026

Related Experiment Video

Updated: May 4, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.0K

An efficient algorithm coupled with synthetic minority over-sampling technique to classify imbalanced PubChem

Ming Hao1, Yanli Wang1, Stephen H Bryant1

  • 1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA.

Analytica Chimica Acta
|December 17, 2013
PubMed
Summary

High-throughput screening often yields imbalanced datasets. A new method combining GLMBoost and Synthetic Minority Over-sampling TEchnique (SMOTE) effectively identifies rare active compounds with improved accuracy and efficiency.

Keywords:
High-throughput screeningImbalanced classificationOver-samplingPubChemUnder-sampling

More Related Videos

Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods
05:34

Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods

Published on: June 6, 2025

1.8K
Efficient Sampling of Genetically Encoded Biosensor Design Space Enabled with a Design of Experiments and Automation Workflow
08:58

Efficient Sampling of Genetically Encoded Biosensor Design Space Enabled with a Design of Experiments and Automation Workflow

Published on: October 17, 2025

840

Related Experiment Videos

Last Updated: May 4, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.0K
Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods
05:34

Applying Cheminformatics to Develop a Structure Searchable Database of Analytical Methods

Published on: June 6, 2025

1.8K
Efficient Sampling of Genetically Encoded Biosensor Design Space Enabled with a Design of Experiments and Automation Workflow
08:58

Efficient Sampling of Genetically Encoded Biosensor Design Space Enabled with a Design of Experiments and Automation Workflow

Published on: October 17, 2025

840

Area of Science:

  • Bioinformatics
  • Machine Learning
  • Computational Chemistry

Background:

  • High-throughput screening (HTS) frequently generates imbalanced datasets.
  • Standard classification methods struggle with imbalanced data, showing poor performance on minority classes.

Purpose of the Study:

  • To develop and evaluate an efficient algorithm for classifying imbalanced datasets from PubChem BioAssay.
  • To improve the detection of rare active compounds in HTS data.

Main Methods:

  • Utilized GLMBoost (Generalized Linear Model boosting) algorithm.
  • Integrated Synthetic Minority Over-sampling TEchnique (SMOTE) to address class imbalance.
  • Compared GLMBoost+SMOTE with Random Forest (RF)+SMOTE on PubChem BioAssay datasets.

Main Results:

  • GLMBoost+SMOTE demonstrated superior performance in classifying rare samples (Sensitivity) and balanced accuracy (Gmean).
  • The proposed combinatorial method significantly improved the detection of active compounds.
  • GLMBoost+SMOTE exhibited greater computational efficiency compared to RF+SMOTE.

Conclusions:

  • The combinatorial approach of GLMBoost and SMOTE is effective for imbalanced classification in HTS.
  • This method offers a promising solution for identifying rare active compounds with high accuracy and efficiency.
  • GLMBoost+SMOTE is recommended for tackling imbalanced classification challenges in cheminformatics and drug discovery.