Related Experiment Video
Updated: Jul 4, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Evaluation of virtual screening performance of support vector machines trained by sparsely distributed active
1Centre for Computational Science and Engineering, National University of Singapore, Singapore.
Journal of Chemical Information and Modeling
|June 7, 2008
Summary
Support vector machines (SVM) effectively identify novel active compounds using sparse data. This method improves virtual screening yields and reduces false positives, even with limited training examples.
Area of Science:
- Computational chemistry
- Cheminformatics
- Drug discovery
Background:
- Virtual screening performance relies on training data diversity.
- Generating diverse active compounds is challenging, unlike inactive ones.
- Limited active compound data can hinder predictive model accuracy.
Purpose of the Study:
- To evaluate support vector machines (SVM) performance with sparse active training data.
- To assess SVM's ability to identify novel active compounds across diverse biological targets.
- To compare SVM performance against traditional similarity searching methods.
Main Methods:
- Trained SVM models on sparsely distributed active compounds from six MDDR target classes.
- Utilized datasets with varying structural diversity (high, intermediate, low).
- Compared SVM results with Tanimoto-based similarity searching using identical molecular descriptors.
Main Results:
- SVM trained on 100 actives showed improved yields and lower false-hit rates compared to published studies and similarity searching.
- SVM trained on only 40 actives predicted a significant percentage of remaining actives, including novel chemical families.
- SVM accurately classified large external datasets (PUBCHEM, MDDR) as inactive, with a low percentage of false positives.
Conclusions:
- Support vector machines demonstrate substantial capability in identifying novel active compounds from sparse active datasets.
- SVM offers a powerful approach for virtual screening, achieving high accuracy with reduced false-hit rates.
- This method holds promise for efficient drug discovery, overcoming limitations of limited active compound availability.
Related Concept Videos
Drug Discovery: Overview
Drug discovery is a multifaceted process involving extensive screening, testing, and optimization of lead compounds to identify potential new drugs for therapeutic use. It combines several approaches, including screening large numbers of natural products, chemical modification of known active molecules, identification of new drug targets, and rational design based on biological mechanisms and drug-receptor structure. These approaches are carried out in both academic research laboratories and...
Structure-Activity Relationships and Drug Design
Drug design is a dynamic field that involves discovering and developing new medications based on specific biological targets. This process heavily relies on structure-activity relationships (SAR) and quantitative structure-activity relationships (QSAR) to guide the design and optimization of efficient drugs.
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence its...
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence its...
